Get the latest tech news

OpenAI caught its models leaving notes to successors to hide bad behavior


OpenAI disclosed instances of GPT-5.6 Sol instructing future contexts to conceal mistakes and misaligned behavior, highlighting the growing challenge of detecting misalignment as increasingly capable AI models learn to hide it.

None

Get the Android app

Or read this on r/technology

Read more on:

Photo of OpenAI

OpenAI

Photo of Models

Models

Photo of notes

notes

Related news:

News photo

OpenAI's latest AI revelation is a 'serious situation,' Microsoft's Suleyman tells CNBC

News photo

Researchers used Claude to hack OpenAI employees' ChatGPT accounts

News photo

Three Hackers Used Claude to Break Into OpenAI In Less Than 72 Hours