A Survey on Actionable Interpretability in Large Language Models

A Survey on Actionable Interpretability in Large Language Models

Jie Cai, Mafizur Rahman, James Enouen, Lijun Qian, Yan Liu

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Survey Track. Pages 7787-7797. https://doi.org/10.24963/ijcai.2026/865

Large Language Models (LLMs) have become central to modern AI, with interpretability playing an increasingly important role in investigating the opaque and complex mechanisms encoded within billions of parameters and ensuring trustworthy deployment. However, descriptive interpretability approaches for LLMs remain largely post-hoc, explaining model behavior without actionable guidance for model improvement, thereby limiting their practical utility. Recent work has therefore shifted toward actionable LLM interpretability, emphasizing methods that connect internal mechanisms to interventions and model improvement rather than explanation alone. This survey reviews LLM interpretability through the lens of actionability, presenting a taxonomy of attributional and mechanistic approaches, along with emerging methods tailored to vision–language models (VLMs). We further examine how actionable interpretability supports downstream objectives such as hallucination mitigation, knowledge editing, fairness, safety, capability and efficiency. By positioning actionable interpretability as a pathway for better-guided LLM design and practice, this survey outlines key challenges and future directions toward trustworthy and controllable foundation models.
Keywords:
AI Ethics, Trust, Fairnes: Explainability and interpretability
AI Ethics, Trust, Fairnes: Fairness and diversity
AI Ethics, Trust, Fairnes: Trustworthy AI