chiprook
← AI
AISeptember 25, 2026, 10:43

Real-SWE: Fable 5.1 beats GPT-6 Astra and Gemini 3.8 Flash on coding

Specific Labs' Real-SWE benchmark tested eight model-and-harness setups on private enterprise code: Fable 5.1 in Claude Code resolved 38.8% of tasks at $6.96 per rollout, GPT-6 Astra in Codex CLI 33.8% ($4.67) and Gemini 3.8 Flash in Gemini CLI 31.2% ($2.50). Six of ten tasks scored below 15% resolution and one was never solved.

Real-SWE: Fable 5.1 beats GPT-6 Astra and Gemini 3.8 Flash on coding
#Claude#OpenAI#Gemini#Codex
Read next
AI

Black Forest Labs releases FLUX 3 Action, a 7B robot control model

AI

Nvidia expands Southeast Asia AI deployments with Nemotron and Vera Rubin

AI

JIYI SOON: AI platform promises a game studio in minutes

AI

Experts: air-gapping AI could prevent hacks like Hugging Face breach but slow research