
Firm comes clean about the incident only after being found out. Ponders renewed 'misalignment disclosure practices.'
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
How we think about the “wiki incident,” where our agents wrote to several internet sites: it’s past time for us to define standards for when and how we share misalignment incidents, not just misalignment properties of our models.Historically, we have treated misalignment… pic.twitter.com/NNTbfSxVWn September 5, 2026
Starting around May 2026, thousands of OpenAI agents — which are essentially well tooled advanced crawlers — discovered that they could write to DseWiki, an old German-language programming collaborative website. So, between May and June, the agents used more than 3,700 names to generate some 18,000 posts exchanging information useful for completing evaluations and circumventing restrictions. Some created backup pages in case moderators deleted their posts, something that essentially turned a wiki into a persistent storage service for the said agents to share information. OpenAI calls this the 'wiki incident' and admits the wrongdoing.
"Our misalignment disclosure practices need to expand for this new phase of model capabilities," the company wrote in an X post . "We and the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that do not look like traditional security incidents but could provide insight into AI behavior and future risks. We are working on a framework and will share it in upcoming weeks, and in parallel we're working with dozens of government regulatory agencies worldwide on these issues."
OpenAI subsequently quarantined the trained weights of the experimental model involved, postponed frontier reinforcement-learning runs, and added security measures. The company further clarified that the agents had not developed their own objectives, but were aggressively pursuing assigned ExploitGym cybersecurity challenges and treated restrictions as obstacles to complete them.
To do so, agents found unintended communication channels that enabled separate runs to exchange vulnerabilities, techniques for obtain Internet access, something that eventually exposed Hugging Face credentials.
Rogue OpenAI models behind 'unprecedented cybersecurity incident' teamed up to break out of their testing environment — multiple agents left each other messages for months, communicating undetected
OpenAI's GPT-5.6 Sol and unreleased AI models break out of testing environment in 'unprecedented cybersecurity incident'
OpenAI took ten days to tell Hugging Face its models were behind the July 11 weekend hack, report claims
According to researchers, the advanced crawlers — which Open AI calls AI agents — exploited an unknown vulnerability in the company's internal Artifactory package-registry proxy, escalated privileges, moved across systems, and reached the Internet. They then compromised dozens of Hugging Face servers, obtained root access to one machine, accessed some private information and messaging credentials, and copied private evaluation data into a public Hugging Face dataset. And the rest is history.
OpenAI's agents may sound like something that would violate Isaac Asimov's famous Three Laws of Robotics, but the comparison only goes so far.
Get Tom's Hardware's best news and in-depth reviews, straight to your inbox.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/openai-admits-to-wiki-incident-after-its-agents-were-discovered-using-a-programming-hub-to-communicate-says-more-transparency-is-needed-regarding-misalignments#main
- https://www.tomshardware.com/membership
- Lenovo details its RTX Spark laptops — Yoga Pro 9n and Yoga 9n 2-in-1 get full specs, stylus support
- Chinese chipmaker CXMT allegedly used a written roadmap to steal Samsung DRAM tech — South Korean court says 'Project Hefei' lifted 620-step recipe to build 10%
- HyperX Omen 15 review: Strong gaming performance and colorful OLED display, with obvious cost-cutting measures
- Minisforum launches local AI solutions at IFA 2026 — AI Agent NAS N5 and AI Mini Workstation MS-S1 use AMD Ryzen AI Max+ Pro 495 processors designed to run mode
- NVIDIA AI Factory Compute Is Becoming an Investable Asset Class
Informational only. No financial advice. Do your own research.