Epoch AI built an IKEA furniture benchmark to test AI models
Nonprofit Epoch AI released the Furniture Assembly Benchmark, which asks AI models to spot errors in partially assembled IKEA furniture using photos and instructions. Anthropic's Claude Opus 4.5 scored just 28%, while OpenAI's GPT-6 Astra reached 80% within 10 months.
- The test uses three IKEA items of varying difficulty: Ställ, Tonstad and Gullaberg
- Models received 60 assembly photos, a PDF manual and a Python interpreter
- Claude Opus 4.5 scored 28%, while GPT-6 Astra reached 80% in 10 months
- GPT-6 Astra takes about 3 minutes per photo, up to 10x faster than rivals
Read next
AI