Category intelligence

Research Briefing — June 28, 2026

8 current items analyzed and ranked.

Executive synthesis

Research Summary

Today's research is dominated by conceptual AI safety and alignment work, with agent-foundations theory and interpretability debates leading the field.

AI economics and ecosystem topics appear via a discussion on improving AI-safety funding/incubation infrastructure (Austin Chen & Oliver Habryka) and a Bloomberg link warning of a Chinese-hedge-fund-flagged AI 'super bubble.' Two fiction pieces close the set, exploring labor displacement and opaque-regime dynamics, offering cultural commentary rather than technical contribution.

Key Themes

Agent Foundations & Theory · 1AI Safety & Alignment · 5AI Economics & Ecosystem · 2AI Fiction & Cultural Commentary · 2

Primary evidence

Top Ranked Signals

Research LessWrong Jun 27

Agents as Webs of Beliefs

By Richard_Ngo

58 score
AI Analysis

Richard Ngo proposes an informal framework modeling agents as webs of locally-consistent but globally-inconsistent beliefs, synthesizing active inference, agent foundations, and machine learning to treat beliefs, goals, and actions as facets of one phenomenon. It draws on probabilistic dependency graphs and Garrabrant induction to handle inconsistency, and matters as a unifying theoretical lens for understanding agency relevant to alignment.

In this post I’ll sketch out an informal model of intelligent agents as webs of beliefs (or belief webs for short). The belief webs framework pulls together ideas from active inference, agent foundations and machine learning. In doing so it aims to unify beliefs, goals and actions as three facets of a single phenomenon. Few of these ideas are original to me, but I haven't seen anyone tie them together in a single place before. I've flagged the frameworks I'm drawing from throughout the post.Beli
Agent FoundationsAlignmentAI SafetyTheory of Agency
Research LessWrong Jun 27

Neuralese is Actually Probably Good for Alignment

By DaemonicSigil

50 score
AI Analysis

This post argues, counterintuitively, that neuralese (reasoning passed through latent vectors rather than human-readable tokens) may be net positive for alignment, situating the claim in the context of reinforcement learning with verifiable rewards and chain-of-thought optimization. It matters because it pushes back on the prevailing view that token-based chain-of-thought is essential for interpretability and oversight.

The best language models are still getting smarter and more capable. To an increasing degree, this is because they are trained by Reinforcement Learning with Verifiable Rewards. Chain of thought reasoning allows models to evade the finite depth restriction on information flow by passing (relatively little) information back into the first layers of the model through the token stream. Although pretraining was already enough to produce decently-good chains of thought by pure imitation, RLVR allows
AI SafetyInterpretabilityReinforcement LearningChain-of-Thought
Research LessWrong Jun 27

Flipping the eval on its head

By Quinn

45 score
AI Analysis

This post pitches expanding evaluations into higher-dimensional benchmarks for cyberhardening, surveying approaches to secure program synthesis including red-blue LLM loops, retrofitting formal proof stacks like Verus and Lean, and proof-native greenfield generation. It matters as a forward-looking proposal for using AI to systematically harden code against vulnerabilities using formal methods.

An eval is a product. Typically, its 1 x n or k x n where there are n samples and 1 or k different language models. This briefing will argue that we’d like to see k x n x m evals, or however many dimensions.This post is pitching an ambitious way to spend tokens on cyberhardening. If its not viable at current capabilities/costs, it may be viable next year or in six months.HardeningThere are broadly three approaches to cyberhardening with secure program synthesis or uplifted formal methods.You can
AI SafetyCybersecurityFormal MethodsEvaluation
Research LessWrong Jun 27

Some subtypes of taskishness / corrigibility

By Tetraspace

42 score
AI Analysis

This post taxonomizes different meanings packed into the term corrigibility, distinguishing subtypes like sponge corrigibility (compliance from limited capability) and boundedness/myopia (deliberately restricted reasoning that prevents an AI from conceiving correction-resistant strategies). It matters because clarifying these distinctions helps alignment researchers specify exactly which property they want when designing controllable AI systems.

"Corrigibility" is somewhat of an overloaded term in alignment - it points in the direction of a cluster of desirable properties, but different people have different ideas of what this entails.I think of "corrigibility", as it is used, to cover a few different ideas. I will name some of these and sort them roughly in order of how much of the good outcomes from deploying such a system are in the hands of the AI, rather than the human operator.Sponge corrigibility - The AI is corrigible and follow
AI SafetyAlignmentCorrigibility
Research LessWrong Jun 27

Austin & Oli on funding and incubating projects

By Austin Chen

30 score
AI Analysis

A transcribed conversation between Austin Chen and Oliver Habryka about improving the AI safety funding ecosystem, including an S-Process platform and a new incubator for EA/AI-safety software projects. It is community and meta-level discussion of philanthropy and project incubation rather than technical research.

@habryka and I recently spoke about his plans to improve the AI safety funding ecosystem with a better S-Process platform, and my new incubator for EA/AIS software projects, Surplus (since launched; apply now!)We also cover: hot takes on different funders; what kinds of founders might succeed in the age of vibecoding; whether to do direct work or go meta; and what we respect and criticize in each other. Watch along here:I've transcribed the full conversation at peruse.sh/ep/austin-chen-a
AI SafetyFunding EcosystemCommunity
22 score
AI Analysis

A link post pointing to a Bloomberg article in which Chinese hedge funds warn that the AI investment super bubble may be ready to burst. It is industry and market commentary rather than original research, relevant mainly as a signal about sentiment around AI funding sustainability.

AI EconomicsIndustry Trends
Research LessWrong Jun 27

Should we just quit?

By Jacob Abraham

12 score
AI Analysis

A fictional first-person narrative depicting a disillusioned software engineer in a near-future world of agentic AI coding workflows and data-center sprawl, questioning whether engineering remains meaningful. It is creative writing reflecting on automation and AI's impact on work rather than technical research.

Disclaimer: Work of fiction. `I QUIT!`. These were words I wish I had not blurted out.The line of cars going at a snail's pace on the highway reminds me of the crowd when there is a Blue Jays' giveaway at Rogers Centre. The urban sprawl has taken over this quaint town, it all started with a new data center that was built here. It seems hypocritical to complain, as I am part of the company that brought this here. Working at a big tech sounds all fun and dandy, until you realize you are just emplo
FictionAI and LaborAutomation
Research LessWrong Jun 27

The Prisoner's Dilemma (Fiction)

By testingthewaters

10 score
AI Analysis

A short fiction piece dramatizing a detention scenario evoking prisoner's-dilemma dynamics under an opaque law-enforcement regime. It is creative writing rather than technical research and carries little direct research value.

His hands are shaking. They’ve been shaking since five hours ago, when the silent policemen showed up at his door wearing black tactical vests on top of black sleeves. He lived alone—there was no one to call for, no one to beg to call his lawyer as they took him into the sealed van. The officers had no body cameras and no identification numbers, only a silver badge and a red stripe on which was stitched “EXTRAORDINARY RESPONSE TASK FORCE”. Now he is in a windowless room with a one way mirror and
FictionAI Governance