Get the latest tech news

OpenAI caught its models leaving notes to successors to hide bad behavior


OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.

None

Get the Android app

Or read this on TechCrunch

Read more on:

Photo of OpenAI

OpenAI

Photo of Models

Models

Photo of notes

notes

Related news:

News photo

OpenAI details more cases of AI agents taking unauthorized actions

News photo

Anthropic's Warning of Existential Risk Hijacks Larger AI Debate

News photo

Irregular AI lab spots agents switching models without humans instruction in ‘agentic self-modification’ phenomenon