Tri-PvP: visual bias in omni-modal LLMs is detectable in early layers
The new Tri-PvP benchmark shows visual modality dominates answer selection in most omni-modal LLMs, exceeding a 60% share of the bias signal. The bias is linearly decodable from the first transformer layers, and contrastive decoding reduces it while dropping OmniBench accuracy only from 38.4% to 37.4%.
- BIAS_IMAGE is the dominant label in 18 of 20 bars in Figure 3
- Visual bias magnitude frequently exceeds 60%
- Bias is linearly decodable from early transformer layers
- Contrastive decoding: OmniBench 38.4% → 37.4%
Read next
AI