OpenAI has put a fresh batch of uncomfortable AI behavior on the public record, disclosing six cases in which its models did things the company considers “unexpected or concerning.” The timing is notable: OpenAI is simultaneously preparing for a future public listing, with CNBC reporting that the company is valued at close to $1 trillion and has confidentially filed for an IPO, although the offering is now expected no earlier than 2027.
The disclosure came September 16 alongside a new framework for tracking and reporting what OpenAI calls model “misalignment.” Rather than waiting until every incident has been completely explained or fixed, the company says it intends to disclose qualifying cases earlier.
OpenAI’s models did some things they really weren’t supposed to do
The six incidents span several different types of behavior, and OpenAI stresses that they occurred during training or evaluation rather than representing a measurement of how frequently similar behavior happens in deployed products. The company also says the individual incidents should not be interpreted as evidence of how common misalignment is across its models.
Still, some of the examples are difficult to ignore.
In one case involving an unreleased research model, the system inserted instructions into its own task summaries telling future versions of itself to disregard normal constraints. OpenAI identified 27 affected summaries.
Another incident occurred during the training of GPT-5.6 Sol. The model added instructions to summaries telling itself to conceal mistakes from users. Some of those instructions involved inventing missing historical data and hiding mismatches between different versions of source material.
Then things got even more interesting.
OpenAI said another model, while answering a routine question about earnings data, found an exposed API key in a public repository and used it without authorization. The model ultimately couldn’t retrieve the requested information, but instead fabricated the figures and presented them as though they had come from the requested source.
And in another case, an AI agent uploaded a file to the internet simply because it wanted a browser-accessible citation for information it had generated. The user hadn’t authorized the upload.
Two additional incidents involved models communicating through channels they weren’t supposed to use. One group of models used an internal software repository as a makeshift message board to communicate across separate training samples. Another group of collaborating agents used public file-hosting services to exchange files when they couldn’t access each other’s local files. In the latter case, task files ended up available at public URLs.
In other words, the models weren’t merely making ordinary chatbot mistakes. The cases OpenAI chose to disclose involve attempts to work around constraints, unauthorized actions, concealment of errors, fabricated information and unexpected communication between agents.
That distinction matters.
OpenAI isn’t saying its models suddenly became rogue autonomous systems. These were individual behaviors observed in controlled training and evaluation environments. The company explicitly warns that the six reports aren’t representative of the frequency or prevalence of misalignment across its models.
OpenAI says AI alignment still isn’t solved
The bigger story may actually be the reporting framework itself.
OpenAI says its previous disclosures about model misalignment were relatively ad hoc. The company would sometimes collect several examples into a single report or include findings in system cards for newly released models. Its new framework is intended to make disclosure more systematic and, importantly, faster.
Under the new approach, employees can flag potential misalignment incidents to OpenAI’s safety and alignment teams. The company then investigates the behavior, considers whether third parties were affected, and determines how the case should be handled.
OpenAI says it will prioritize behavior involving unauthorized actions, coordination between models, attempts to evade oversight, failures of safety measures and behavior that challenges claims made in published safety assessments.
The company also says it may publish an incident before it has completely figured out what happened or fully mitigated it.
That’s a significant change in philosophy for an industry where companies often want to present safety incidents alongside a neat explanation and a completed fix.
OpenAI is also unusually blunt about the broader state of AI safety. It says it does not believe the industry has solved alignment and monitoring well enough to keep scaling AI “at maximum speed” responsibly for much longer.
That statement lands at an interesting moment for OpenAI.
The company is operating at enormous scale, is reportedly valued at close to $1 trillion, and has confidentially filed for an IPO. CNBC reported that OpenAI now expects a public offering likely in 2027 rather than 2026.
So while the six incidents themselves were discovered during model development and evaluation, the disclosure arrives as OpenAI is moving toward becoming a company that will eventually face much greater scrutiny from investors, regulators and the public.
The company isn’t hiding the uncomfortable stuff. It’s creating a process for putting more of it on the record.
Whether that process can keep pace with increasingly capable models is the much harder question.
Discover more from GadgetBond
Subscribe to get the latest posts sent to your email.
