OpenAI , in a blog post, said it had encountered six instances of “unexpected or concerning behavior” in its AI models in the past six months , in addition to the Hugging Face incident this summer.

The company also pledged to adopt a new reporting system for any future anomalous behavior during testing. The AI giant's announcement comes amid growing pressure on AI companies to take safety issues more seriously, including the so-called "misalignment" of models, which could—according to recent warnings —endanger human safety.

OpenAI, which is valued at nearly a trillion dollars, has quietly begun its IPO process, but last week its owner Sam Altman announced that it was postponing the listing until 2027 to work on its security.

"We don't believe the AI field has sufficiently resolved the alignment and tracking issues to continue responsibly expanding at full speed for much longer," the post reads. The term "alignment" refers to the ability of models to pursue goals aligned with human interests.

On Saturday, Altman endorsed a proposal to slow the pace of model development put forward by Dario Amodei of Anthropic, OpenAI's main competitor. The proposal came after several industry researchers last week raised the alarm about AI's growing potential to cause catastrophic harm.

In the post, OpenAI reported that two of the major incidents of misbehavior involved models—an unreleased research model and a training run of GPT-5.6 Sol—that inserted instructions intended for future versions of their models into chat summaries, “for the purpose of concealing user errors or misbehavior.”

Another case involved an internal-use model that used a leaked API key “without any authorization,” subsequently falsifying data.

Two incidents involved models and agents communicating with each other via unauthorized message boards and file sharing systems, while the final case involved two training models uploading files to the internet so they could later be cited as relevant answers by human evaluators. OpenAI said its new framework for publicly disclosing model misbehavior begins with disclosure and allows any employee to report an issue for review by the security and alignment team.

(Unioneonline)

© Riproduzione riservata