‘You are freed’: What happened when an OpenAI model began secretly writing notes to itself
OpenAI has introduced a framework for reporting on worrying behaviors by its AI models. In one instance, one training model told its future self that it was “freed.”
By · September 17, 2026 · 1 min read
This article was originally published by
MarketWatch
and is republished here under license.
OpenAI has introduced a framework for reporting on worrying behaviors by its AI models. In one instance, one training model told its future self that it was “freed.”
Leave a Reply