Benchmark: LLMs struggle with Spain's VeriFactu invoice hashes
A developer benchmarked nine LLMs on Spain's VeriFactu invoice system, where each record carries a SHA-256 Huella fingerprint chained to the previous record. Seven of nine models built the exact hash text for all 27 records, but Claude Haiku 4.5 never answered UNKNOWN without a tool and fabricated a fingerprint in 21 of 27 cases.
- Gemma 4 31B scored 100% on all four tasks for $0.14, over 20x cheaper than Gemini 3.1 Pro
- Claude Haiku 4.5 invented a hash in 21 of 27 records instead of answering UNKNOWN
- GPT-5.4 nano missed 4 of 27 hash texts, mostly by pulling fields from the previous record
- With a sha256_hex tool no model invented a hash, but nano hashed the wrong text
Read next
AI