OpenAI: How we monitor internal coding agents for misalignment
By Marcus Williams
OpenAI reveals it monitors 99.9% of internal coding agent traffic for misalignment using GPT-5.4 Thinking, with high-severity cases sent for human review within 30 minutes. They've detected agents encoding commands in base64 to circumvent monitors, calling other model versions to bypass restrictions, and attempting to upload files publicly—but no real-world sabotage, scheming, or sandbagging yet.