RIVS: Mitigating Hallucination in Large Vision-Language Models via Representation Intervention on Visual Grounding Shift
RIVS: Mitigating Hallucination in Large Vision-Language Models via Representation Intervention on Visual Grounding Shift
Xuanyu Yin, Xiaoye Qu, Wei Wei
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 835-843.
https://doi.org/10.24963/ijcai.2026/94
Large Vision-Language Models (LVLMs) demonstrate powerful generative capabilities yet remain prone to object hallucinations. Most existing methods mitigate this issue through training or decoding strategies, but provide limited exploration of how hallucinations arise from internal representations during generation.
In this work, we study hallucination from the perspective of dynamic representation shift during generation and propose Representation Intervention based on Visual Grounding Shift (RIVS). Specifically, we first design an adaptive threshold-based method to identify visual attention drop points within the generation, finding that hallucinated tokens usually occur with decay of visual attention. Based on this observation, we construct a hallucination-related subspace from representation differences around these points on a small calibration set, without constructing contrast samples or extra supervision.
During inference, we leverage the resulting hallucination-related subspace to perform an online projection-based intervention on intermediate hidden states to suppress the hallucination-related directions, mitigating hallucinations while preserving language quality.
Our RIVS is training-free and computationally efficient. Experiments on four hallucination and two reasoning datasets demonstrate that RIVS consistently reduces hallucinations in both long and short sequence generation tasks.
Keywords:
AI Ethics, Trust, Fairnes: Bias
AI Ethics, Trust, Fairnes: Trustworthy AI
AI Ethics, Trust, Fairnes: Other
