chiprook
← AI
AIOctober 6, 2026, 21:31

Benchmark tests whether LLMs invent Japanese law articles

A developer checked Japanese law article citations from three local LLMs (1.8B–8B parameters) against the official e-Gov registry of 6,913 articles. llm-jp-3-1.8b invented 4.57% of articles in Japanese answers versus 1.09% in English, Qwen2.5-7B 1.17% versus 0%, and Swallow-8B none in either language.

Benchmark tests whether LLMs invent Japanese law articles
#Qwen#Swallow#Llm-jp
Read next
AI

Open-source benchmark tests whether AI agents can engineer working robots

AI

Benchmark: LLMs struggle with Spain's VeriFactu invoice hashes

AI

Benchmark: top LLMs know only 24% of post-cutoff 2025-2026 facts

AI

Benchmark: flagship LLMs cave to user pressure more often