Mutual Information Guided Reinforcement Learning for Ambiguous Label Disambiguation

Mutual Information Guided Reinforcement Learning for Ambiguous Label Disambiguation

Jinfu Fan, Jiangnan LI, Xiaohui Zhong, Kangrui Ren, Linqing Huang

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 4260-4268. https://doi.org/10.24963/ijcai.2026/474

Partial-label learning (PLL) addresses challenging scenarios where each instance is associated with a set of candidate labels and only one is the truth. Most existing PLL methods rely on static disambiguation heuristics, which are prone to error propagation when the ambiguity labels are high. To address this issue, we propose a novel mutual information guided reinforcement learning framework for partial label disambiguation (MR-PLL). In this framework, label disambiguation is formulated as a sequential Markov decision process, where an agent dynamically discriminates whether to retain, modify, or abstain from correcting ambiguous labels. To ensure reliable decision making, we introduce a mutual information guided gating mechanism that adaptively adjusts the confidence of soft label propagation according to the dependency between feature representations and labels. The abstention mechanism allows the model to postpone uncertain decisions, resulting in more reliable disambiguation. Furthermore, we design an information weighted reward term during the actor-critic process to gradually improve the label disambiguation ability of the policy, and provide theoretical analysis on convergence and bias reduction. Experiments on benchmark and real datasets verify the effectiveness of the proposed algorithm.
Keywords:
Machine Learning: Classification
Machine Learning: Clustering
Machine Learning: Multi-label learning
Machine Learning: Weakly supervised learning