Real-SWE: Fable 5.1 beats GPT-6 Astra and Gemini 3.8 Flash on coding
Specific Labs' Real-SWE benchmark tested eight model-and-harness setups on private enterprise code: Fable 5.1 in Claude Code resolved 38.8% of tasks at $6.96 per rollout, GPT-6 Astra in Codex CLI 33.8% ($4.67) and Gemini 3.8 Flash in Gemini CLI 31.2% ($2.50). Six of ten tasks scored below 15% resolution and one was never solved.
- Fable 5.1 via Claude Code: 38.8% resolution at $6.96 per rollout
- GPT-6 Astra via Codex CLI: 33.8% at $4.67; Gemini 3.8 Flash: 31.2% at $2.50
- Six of ten tasks scored below 15% resolution, one was never solved
- Missed requirements, not broken syntax, dominate failures
Read next
AI