LiveNerf publicly tracks Claude Opus 5.5 degradation
Since September 24, 2026, the LiveNerf project has been measuring Claude Opus 5.5 daily on a fixed panel of 78 questions, comparing results against the model's launch baseline. The method pins prompts, CLI version and harness hash, and tracks token usage as an early signal of reduced reasoning effort.
- The 78-question panel was filtered from 2,336 candidates by unstable answers
- Lower reasoning effort cuts output tokens by 62% while accuracy drops 8.3 points
- A daily panel run detects about a 7.5-point accuracy shift per 10-day window
- The project uses the Inspect framework and Anthropic's paired-comparison statistics
Read next
AI