OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models’ hidden reasoning

OpenAI says actors linked to China-based Moonshot AI spearheaded a campaign to extract its models’ hidden reasoning

When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .

Activity began on July 1st, “initially at a low volume.” After that came “high-volume spikes” on July 24 and 25, with 16,000 requests using an extraction pattern. The requests came from over 4,000 users. The activity attempted extraction but was “not necessarily successful,” and the campaign was “fully disrupted by July 28.” The post did not clarify any measure of success rate, which models were specifically targeted, or how many of the users were Moonshot-linked.

OpenAI defines protected reasoning as “the model’s internal record for working through a task” and adversarial distillation as the “systematic and unauthorized use of one model’s outputs or reasoning” to train or improve another model. This data is encrypted to hide the model’s chain of thought and is handed to the client as an encrypted block. The client sends that block back with each request, so the provider doesn’t have to store it. One method the operators tried took the encrypted reasoning from one conversation and asked a model in another to decrypt it.

OpenAI says the encryption, in this case, was not broken. There was no direct access to stored user conversations, and no database was compromised. One fix closed a pathway that let someone who already had another user’s encrypted reasoning replay it and recover its contents. Separately, it added checks to detect and hold streamed output that might expose reasoning. It also strengthened protections for hidden reasoning across users, workspaces, organizations, and model families, and worked with third-party providers to disrupt accounts whose activity moved through their services.

Independent security researchers had also brought “related cross-model and conversation-compaction vulnerabilities” to OpenAI through responsible disclosure, and the company confirmed the attack paths they found were real. The paper “Stealing Reasoning Traces from Proprietary LLM APIs,” dated Aug. 10, is explicitly linked by the post. Testing OpenAI, Anthropic, and Google , the researchers fed a frontier model’s encrypted reasoning to a corresponding weaker model, which then wrote it out in plain text. The researchers ran their test in early July and, after the providers acknowledged their report, they were “unable to launch the same attacks.”

OpenAI took ten days to tell Hugging Face its models were behind the July 11 weekend hack, report claims

OpenAI agent goes rogue and hacks popular AI community

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment