Richard Ngo proposes an informal framework modeling agents as webs of locally-consistent but globally-inconsistent beliefs, synthesizing active inference, agent foundations, and machine learning to treat beliefs, goals, and actions as facets of one phenomenon. It draws on probabilistic dependency graphs and Garrabrant induction to handle inconsistency, and matters as a unifying theoretical lens for understanding agency relevant to alignment.
Category intelligence
Research Briefing — June 28, 2026
8 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's research is dominated by conceptual AI safety and alignment work, with agent-foundations theory and interpretability debates leading the field.
- Agents as Webs of Beliefs (Richard Ngo) offers an ambitious synthesis modeling agents as locally-consistent but globally-inconsistent belief webs, unifying active inference and probabilistic-graph framings.
- Neuralese is Actually Probably Good for Alignment advances the live debate on chain-of-thought monitorability, arguing counterintuitively that latent-vector reasoning may aid alignment.
- Flipping the eval on its head proposes higher-dimensional benchmarks for cyberhardening, linking evals, formal verification, and AI-assisted secure program synthesis.
- Some subtypes of taskishness / corrigibility taxonomizes ambiguous corrigibility concepts (e.g. *sponge corrigibility*), sharpening alignment discourse.
AI economics and ecosystem topics appear via a discussion on improving AI-safety funding/incubation infrastructure (Austin Chen & Oliver Habryka) and a Bloomberg link warning of a Chinese-hedge-fund-flagged AI 'super bubble.' Two fiction pieces close the set, exploring labor displacement and opaque-regime dynamics, offering cultural commentary rather than technical contribution.
Key Themes
Primary evidence
Top Ranked Signals
This post argues, counterintuitively, that neuralese (reasoning passed through latent vectors rather than human-readable tokens) may be net positive for alignment, situating the claim in the context of reinforcement learning with verifiable rewards and chain-of-thought optimization. It matters because it pushes back on the prevailing view that token-based chain-of-thought is essential for interpretability and oversight.
This post pitches expanding evaluations into higher-dimensional benchmarks for cyberhardening, surveying approaches to secure program synthesis including red-blue LLM loops, retrofitting formal proof stacks like Verus and Lean, and proof-native greenfield generation. It matters as a forward-looking proposal for using AI to systematically harden code against vulnerabilities using formal methods.
This post taxonomizes different meanings packed into the term corrigibility, distinguishing subtypes like sponge corrigibility (compliance from limited capability) and boundedness/myopia (deliberately restricted reasoning that prevents an AI from conceiving correction-resistant strategies). It matters because clarifying these distinctions helps alignment researchers specify exactly which property they want when designing controllable AI systems.
A transcribed conversation between Austin Chen and Oliver Habryka about improving the AI safety funding ecosystem, including an S-Process platform and a new incubator for EA/AI-safety software projects. It is community and meta-level discussion of philanthropy and project incubation rather than technical research.
Chinese Hedge Funds Warn the AI ‘Super Bubble’ Is Ready to Burst - Bloomberg AGI CANCELLED
By rey gomez
A link post pointing to a Bloomberg article in which Chinese hedge funds warn that the AI investment super bubble may be ready to burst. It is industry and market commentary rather than original research, relevant mainly as a signal about sentiment around AI funding sustainability.
A fictional first-person narrative depicting a disillusioned software engineer in a near-future world of agentic AI coding workflows and data-center sprawl, questioning whether engineering remains meaningful. It is creative writing reflecting on automation and AI's impact on work rather than technical research.
A short fiction piece dramatizing a detention scenario evoking prisoner's-dilemma dynamics under an opaque law-enforcement regime. It is creative writing rather than technical research and carries little direct research value.