chiprook
← AI
AIOctober 6, 2026, 21:15

Benchmark: LLMs struggle with Spain's VeriFactu invoice hashes

A developer benchmarked nine LLMs on Spain's VeriFactu invoice system, where each record carries a SHA-256 Huella fingerprint chained to the previous record. Seven of nine models built the exact hash text for all 27 records, but Claude Haiku 4.5 never answered UNKNOWN without a tool and fabricated a fingerprint in 21 of 27 cases.

Benchmark: LLMs struggle with Spain's VeriFactu invoice hashes
#Anthropic#Google#OpenAI#Gemma
Read next
AI

Benchmark: top LLMs know only 24% of post-cutoff 2025-2026 facts

AI

Benchmark: flagship LLMs cave to user pressure more often

AI

New Benchmark Tests 24 LLMs Against Human Writers on 475 Prompts

AI

Preprint Finds a Steerable 'Pain Axis' in 25 Open-Weight LLMs, Not Claude