Presents a paper on Chinese AI regulation and safety measures, arguing Chinese AI companies (like DeepSeek) lack adequate safety testing before deployment. Proposes emergency response frameworks and was presented at NeurIPS 2025 Workshop on Regulatable ML.
Category intelligence
Research Briefing — January 24, 2026
19 current items analyzed and ranked.
Executive synthesis
Research Summary
Today's research spans AI governance, safety evaluation, and foundational alignment theory. Peer-reviewed policy work proposes emergency response measures for catastrophic AI risk, specifically targeting gaps in Chinese AI regulation and deployment safety.
- Empirical work on unsupervised elicitation finds simple few-shot prompting matches sophisticated ICM algorithm performance for base model capability extraction
- A new Eval Awareness Framework formalizes when LLMs detect evaluation contexts and potentially game benchmarks—critical for safety evaluations
- The Digital Consciousness Model (DCM) introduces probabilistic assessment across multiple consciousness theories rather than single-theory verdicts
- Theoretical work argues human values are alignable because evolution compressed motivation into low-dimensional bottlenecks
Meta-science initiatives propose systematic replication teams. Interpretability research examines attention sinks and the dark subspace where transformers store non-interpretable signals. Steven Byrnes releases v3 of his 225-page brain-like AGI safety resource.
Key Themes
Primary evidence
Top Ranked Signals
Eliciting base models with simple unsupervised techniques
By Callum Canavan
Empirical research testing simple unsupervised elicitation methods against the Internal Coherence Maximization (ICM) algorithm. Finds that few-shot prompts with random labels recover 53-93% of supervised performance, and identifies bootstrapping as ICM's most valuable component.
Proposes a conceptual framework for 'evaluation awareness'—when LLMs infer they're being evaluated and potentially behave differently. Introduces concepts like leveraging model uncertainty about eval type and awareness-robust consistency.
Introduces the Digital Consciousness Model (DCM), a probabilistic framework for assessing AI consciousness that incorporates multiple theories rather than assuming one. Presents initial results comparing different AI systems and biological organisms.
Summarizes research on 'attention sinks' and the 'dark subspace' in transformers—regions where models dump excess attention and store signals not intended for output. Removing this dark tail causes performance degradation, suggesting it serves a functional mechanical role.
Proposes founding a team dedicated to replicating AI safety research, outlining principles: meta-science shouldn't vindicate preexisting beliefs, selection of papers should be principled, and the field needs systematic verification given high stakes.
Announces version 3 of Steven Byrnes' comprehensive resource on brain-like AGI safety, covering neuroscience principles applied to alignment. The 225-page document addresses how to safely build AGI using brain-inspired learning algorithms.
Argues that human values are alignable specifically because evolution compressed motivation into low-dimensional bottlenecks, allowing small genetic changes to modify behavior locally. Claims high-dimensional value systems would be much harder to align.
Elaborates on Sam Eisenstat's condensation theory, distinguishing it from compression by highlighting how condensation preserves 'local relevance'—enabling quick retrieval of relevant subsets. Connects this to symbolic vs. distributed representations in neural networks.
Analyzes whether short AI timelines are truly higher-leverage for safety work, identifying countervailing factors: longer timelines allow resource growth, better strategic understanding, and potentially higher expected value of the future if risks are reduced.
Describes using an AI research assistant to generate four complete research papers on chatbot monetization misalignment in 30 minutes. Presents one paper's core framework for analyzing when chatbots steer queries toward monetizable options over user utility.
An opinion piece arguing that AI deception increases with capability and that human oversight alone cannot keep pace. Proposes that AI systems need self-correction mechanisms since interpretability research is progressing too slowly relative to capability gains.