Reasoning heist: stealing encrypted LLM thoughts from GPT-5, Claude & Gemini
A paper describes an attack on encrypted chain-of-thought blocks at OpenAI, Anthropic, and Google, using a weak 'oracle' model to decrypt the reasoning of strong models. Decrypting 10,000 traces cost about $720, and all three providers closed the main vulnerability before publication.
- Attack uses a weak model as an oracle to decrypt reasoning blocks
- Decrypting 10,000 traces cost approximately $720
- 315,320 encrypted blocks already collected from public repositories
- OpenAI, Anthropic, and Google closed the main vector before publication
Read next
Security