Chat template triggers "I'm just an AI" disclaimer in 8 LLMs
An August 9, 2026 paper shows the phrase "I'm just a language model" depends not on model weights but on the chat template: with it, models give the standard disclaimer; without it, the same model speaks in first person. The effect held across 8 open-source instruction models up to 9B parameters, and in 3 of them a steering direction in activations was found to control the behavior.
- 8 open-source instruction models up to 9B parameters were tested, the effect repeated in all
- With the chat template the disclaimer rises; without it, first-person answers with "feel" and "think"
- In 3 models a steering direction was found: adding it triggers the disclaimer, subtracting it disables it
- A random vector of the same norm had no effect; accepted at a COLM 2026 workshop and KONVENS 2026 Eval4SD
Read next
AI