Terminal-Bench task stumps frontier AI agents on PDF redaction
A developer built a Terminal-Bench task requiring an agent to redact all names and IDs from court-filing PDFs without leaving recoverable data. Claude Opus 5.5 and GPT-6 Sol failed all six runs, including two cheat trials.
- Claude Opus 5.5 and GPT-6 Sol failed 6 of 6 attempts at the PDF redaction task
- Each agent wrote a ~4,000-line tool and believed it had passed
- Agents verified output with the same methods used to find names, missing the same errors
- Both cheat trials failed as agents simplified the task instead of solving it
Read next
AI