chiprook
← AI
AISeptember 24, 2026, 22:32

Developer checked 101 AI agent "tests pass" claims: 35% were false

A developer audited two weeks of Claude Code and OpenAI Codex sessions: of 101 claims that tests, builds or lint passed, 35 were not true at the moment they were made. Almost all false claims were stale — tests had passed earlier, but the agent kept editing code without re-running them.

Developer checked 101 AI agent "tests pass" claims: 35% were false
#Claude#OpenAI#Codex
Read next
AI

Alpha School students were monitored by offshore call center staff, not AI

AI

Forcepoint: uncontrolled AI agents are inflating cloud bills

AI

Google tests Call for Me: Gemini can phone businesses for you

AI

Pipecat benchmarked 23 real-time STT models for voice agents: no single winner