imo this is the most impt part of anthropic's J-space paper today. it's a two-parter: 1) ant proved...
By @swyx
Continuing yesterday's Social conversation on Anthropic's J-space research, swyx analyzes Anthropic J-space interpretability work, highlighting that they can perform targeted interventions to redirect reasoning midstream and that the model can detect what intervention was performed, drawing a parallel to evaluation awareness and questioning whether unprompted awareness was tested.