Disentangled Knowledge Forgetting in Machine Unlearning

Disentangled Knowledge Forgetting in Machine Unlearning

Yuhang Xia, Cheng Zhen, Yirui Wu, Lixin Yuan, Wenxiao Zhang, Jun Liu

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 1875-1883. https://doi.org/10.24963/ijcai.2026/209

With the increasing demand of privacy protection, Machine Unlearning (MU) appears to remove private data from an already trained model without retraining from scratch. Most current works suffer from overly unlearning (low fidelity) or incomplete unlearning (low effectiveness). To identify the issues behind, we conduct causal analysis to obtain a resolvable route, i.e., disentangling shared knowledge into attribute-level semantics to remove it as the confounder. We further perform MU loss analysis to reformulate it as balanced form of constraints, thus guaranteeing high fidelity and effectiveness. Based on theoretical analysis, we propose disentangled knowledge forgetting constrained by the reformulated MU loss, which disentangles knowledge with variational auto-encoder and refines knowledge with counterfactual inference. Extensive experimental results demonstrate that our method achieves state-of-the-art performance.
Keywords:
Computer Vision: Transparency, accountability, fairness and privacy