chiprook
← AI
AIOctober 8, 2026, 00:46

Terminal-Bench task stumps frontier AI agents on PDF redaction

A developer built a Terminal-Bench task requiring an agent to redact all names and IDs from court-filing PDFs without leaving recoverable data. Claude Opus 5.5 and GPT-6 Sol failed all six runs, including two cheat trials.

Terminal-Bench task stumps frontier AI agents on PDF redaction
#Anthropic#OpenAI#Claude#GPT
Read next
AI

Anthropic releases Claude Sonnet 5.5: 70.6% on Terminal-Bench 4.0 at $2/$10

AI

SpaceXAI launches Grok 4.7: half the price of rivals, Terminal-Bench nearly doubles

Business

Botched PDF Redactions Expose Google Nebraska Data Centers' Water and Power Use

AI

Scientists Loaded Frontier AI Models Into a Self-Driving Toyota Corolla; Only GPT-6 Astra Finished the Course