The Encryption Was the Target: How a 'Fuzzy Decoder' Trick Turned Reasoning Protection Into an Attack Surface
OpenAI encrypts the chain-of-reasoning tokens its models return through the API specifically so that competitors cannot read how a model thinks, only its final answer. The operators OpenAI disrupted found a way around that: they copied a model's encrypted reasoning output from one conversation, pasted it into an entirely separate conversation, and asked a different instance of the model to decrypt and transcribe it in plain text [1]. OpenAI described this as manipulating model interactions so protected reasoning could be reproduced in forms visible to the requester, carried out in a coordinated, scaled manner that violated its terms of service [2].
The trick worked because of a deeper architectural problem, not just a prompt-engineering loophole. An academic paper on 'Stealing Reasoning Traces from Proprietary LLM APIs' found that encrypted reasoning blocks from OpenAI, Anthropic, and Google were interchangeable across sessions, users, and even different models, effectively letting a weaker model act as a fuzzy decoder of a stronger model's hidden reasoning [4]. The implications went beyond model copying: researchers who decoded 315,320 thinking blocks drawn from 6,708 public agent trajectories recovered 704 distinct privacy artifacts from genuine user sessions, including 62 API keys, 33 passwords, 24 access tokens, and 7 private keys [4]. What was marketed as a protected, hidden layer of the model's cognition turned out to be a container anyone with the right replay trick could open.


