Category intelligence

Research Briefing — January 3, 2026

14 current items analyzed and ranked.

Executive synthesis

Research Summary

This batch features notable technical contributions alongside alignment theory and community analysis. Instruct Vectors demonstrates that steering vectors trained on frozen base models can induce consistent assistant behavior without traditional post-training, offering new insight into what instruction-tuning actually accomplishes.

Supplementary work includes 2025 prediction calibration (finding forecasts were overestimated), educational alignment content covering distributional leap problems, and speculative explorations drawing on developmental psychology and evolutionary biology for alignment insights.

Key Themes

Language Models · 3AI Safety/Alignment · 7AI Forecasting · 1Off-topic/Non-AI · 4

Primary evidence

Top Ranked Signals

72 score
AI Analysis
Demonstrates that steering vectors trained on frozen base models can induce consistent assistant behavior without traditional post-training. Qwen3-4B-Base successfully imitates instruction-tuned behavior using per-layer vectors.
Post-training is not necessary for consistent assistant behavior from base modelsImage by Nano Banana ProBy training per-layer steering vectors via descent on a frozen base model, I found that it is possible to induce consistent assistant behavior, including the proper use of EOS tokens at the end of assistant turns and consistent reference to the self as an AI assistant. Using the steering vectors, Qwen3-4B-Base was able to imitate the behavior of an instruction/chat tuned model.Many of the ima
Language ModelsActivation SteeringMechanistic InterpretabilityAlignment
Research LessWrong Jan 1

Debunking claims about subquadratic attention

By Vladimir Ivanov

70 score
AI Analysis
Critical analysis arguing that claimed subquadratic attention mechanisms (Kimi Linear, DeepSeek Sparse Attention, Mamba, RWKV) either remain quadratic in practice or underperform standard attention on capability benchmarks.
TL;DR: In the last couple years, there have been multiple hype moments of the form "<insert paper> figured out subquadratic/linear attention, this is a game changer!" However, all the subquadratic attention mechanisms I'm aware of either are quadratic the way they are implemented in practice (with efficiency improved by only a constant factor) or underperform quadratic attention on downstream capability benchmarks. A central issue with attention is that its FLOP complexity is qua
Language ModelsTransformer ArchitectureEfficiencyTechnical Analysis
Research LessWrong Jan 2

Scale-Free Goodness

By testingthewaters

55 score
AI Analysis
Proposes 'scale-free alignment' where aligned AI behavior remains understandable and approvable by less intelligent actors even as AI capabilities increase. Argues good actors should be 'good-registering' across intelligence scales.
Introduction Previously I wrote about what it would mean for AI to “go well”. I would like to elaborate on this and propose some details towards a “scale-free” definition of alignment. Here “scale-free alignment” means a version of alignment that does not feature sudden and rapid “phase shifts”, so as aligned actors get more intelligent their behaviour remains understandable and approved by less intelligent actors. In other words, there should be no moment where a superintelligence looks at us a
AI SafetyAlignmentSuperintelligence
Research LessWrong Jan 2

Where do AI Safety Fellows go? Analyzing a dataset of 600+ alumni

By Christopher_Clay

55 score
AI Analysis
Empirical analysis of 600+ alumni from 9 major AI safety fellowships, tracking career outcomes. Finds 10%+ of fellows did another fellowship afterward, questioning efficiency of current pipeline.
We invest heavily in fellowships, but do we know exactly where people go and the impact the fellowships have? To begin answering this question I manually analyzed over 600 alumni profiles from 9 major late-stage fellowships (fellowships that I believe could lead directly into a job following). These profiles represent current participants and alumni from MATS, GovAI, ERA, Pivotal, Talos Network, Tarbell, Apart Labs, IAPS, and PIBBS.Executive SummaryI’ve compiled a dataset of over 600 alumni prof
AI SafetyCommunity BuildingCareer Trajectories
Research LessWrong Jan 1

2025 in AI predictions

By jessicata

50 score
AI Analysis
Annual evaluation of AI predictions, finding 2025 predictions mostly overestimated capabilities. Notes 'AGI' is becoming less useful as a term and identifies cluster of predictions expecting large AI effects by 2030.
Past years: 2023 2024Continuing a yearly tradition, I evaluate AI predictions from past years, and collect a convenience sample of AI predictions made this year. I prefer selecting specific predictions, especially ones made about the near term, enabling faster evaluation.Evaluated predictions made about 2025 in 2023, 2024, or 2025 mostly overestimate AI capabilities advances, although there's of course a selection effect (people making notable predictions about the near-term are more l
AI ForecastingPredictionsAGI Timelines
45 score
AI Analysis
Educational post in an alignment series covering the 'distributional leap' problem, how critics learn to predict reward, and four key difficulties in ensuring AI systems learn desired values.
2.1 SummaryIn the last post, I introduced model-based RL, which is the frame we will use to analyze the alignment problem, and we learned that the critic is trained to predict reward.I already briefly mentioned that the alignment problem is centrally about making the critic assign high value to outcomes we like and low value to outcomes we don’t like. In this post, we’re going to try to get some intuition for what values a critic may learn, and thereby also learn about some key difficulties of t
AI SafetyAlignmentReinforcement LearningEducation
Research LessWrong Jan 2

On Moral Scaling Laws

By unduePestilence

40 score
AI Analysis
Explores how moral weight might scale with the mental complexity of moral patients, examining implications for utilitarian ethics and effective altruism cause prioritization (e.g., shrimp welfare).
INTRODUCTIONIn Utilitarian ethics, one important factor in making moral decisions is the relative moral weight of all moral patients affected by the decision. For instance, when EAs try to determine whether or not shrimp or bee welfare (or even that of chickens or hogs) is a cause worth putting money and effort into advancing, the importance of an individual bee or shrimp’s hedonic state (relative that of a human, or a fish, or a far-future mind affected by the long-term fate of civilization) is
AI EthicsPhilosophyEffective Altruism
35 score
AI Analysis
Speculative exploration of whether community-based RLHF might inadvertently lead AI systems to adopt social consensus as a proxy for truth. Raises concerns about social epistemology being embedded in AI alignment approaches.
tl;dr: rambling thoughts on why community-based RLHF might help with alignment but have the unintended bad consequence of effectively adopting a social epistemology/consensus theory of truth in general. Epistemic status: this is Part 3[1] of a raw and unfiltered brain dump of the notes I jotted down while attending NeurIPS and its adjacent workshops in December. None of it has been thought through deeply, it's not carefully written and there are no pretty pictures. But I won’t hav
AI SafetyAlignmentRLHF
30 score
AI Analysis
Brainstorms whether insights from developmental cognitive psychology (Piaget, Montessori) could inform curriculum design for AI training to produce more aligned, interpretable models.
tl;dr: brainstorming on alternative curriculum approaches to training models that might cause the embedded knowledge to be structured in a better way (more truth aligned, interpretable, corrigible...)Epistemic status: this is Part 2[1] of a raw and unfiltered brain dump of the notes I jotted down while attending NeurIPS and its adjacent workshops in December. None of it has been thought through deeply, it's not carefully written and there are no pretty pictures. But I won’t have time to res
AI SafetyAlignmentTraining Methods
30 score
AI Analysis
Brainstorms potential analogies between biological evolution mechanisms and AI model development, seeking ideas that might help alignment research. Considers incrementalism and selection pressures.
tl;dr: brainstorming about mechanisms that operate in biological evolution, and wondering whether/how analogous evolutionary mechanisms now exist for AI models, or could, or should.  Undertaken with the broad motivation of mining biology for ideas that might help alignment research.Epistemic status: this is part 1[1] of a raw and unfiltered brain dump of the notes I jotted down while attending NeurIPS and its adjacent workshops in December.  None of it has been thought through dee
AI SafetyAlignmentEvolutionary Computation
Research LessWrong Jan 2

2025 Letter

By zef

10 score
AI Analysis
Personal reflective essay about 2025, blending poetry with thoughts on AI acceleration and existential uncertainty. Primarily emotional/philosophical rather than analytical.
I wrote a letter this year about 2025. It's about acceleration, poetry, how it's been the most eventful year of my life, and how I am excited and scared for the future. Crossposted from my substack.Letter I want to tell you a story about 2025. As I bump along today and approach 21 on into the new year, in a van riding from Burgundy to Paris, and I stare at the small hills, the snow inscribed against the mud like frosted chocolate, extending down into the highway and then melting over into t
Personal ReflectionOff-topic
Research LessWrong Jan 2

The Weirdness of Dating/Mating: Deep Nonconsent Preference

By johnswentworth

5 score
AI Analysis
A social psychology post exploring theories about nonconsent preferences in dating/mating contexts, citing survey statistics. Not related to AI research or technology.
Every time I see someone mention statistics on nonconsent kink online, someone else is surprised by how common it is. So let’s start with some statistics from Lehmiller[1]: roughly two thirds of women and half of men have some fantasy of being raped. A lot of these are more of a rapeplay fantasy than an actual rape fantasy, but for purposes of this post we don’t need to get into those particular weeds. The important point is: the appeal of nonconsent is the baseline, not the exception, especiall
Off-topicSocial Psychology