Measuring the ability of Opus 4.5 to fool narrow classifiers
By Fabien Roger
Researchers measure Claude Opus 4.5's ability to generate adversarial attacks that fool prompted and fine-tuned classifiers (monitors) in a security-relevant context (BashBench). They find relatively low attack success rates, especially against chain-of-thought classifiers and fine-tuned Haiku 4.5, and present a methodology for evaluating future AI monitors.