Toward Trustworthy Foundation Models: Mitigating OOD Failure in Adaptation and Deployment

Toward Trustworthy Foundation Models: Mitigating OOD Failure in Adaptation and Deployment

Ping Song

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Doctoral Consortium. Pages 8339-8340. https://doi.org/10.24963/ijcai.2026/948

Foundation models’ versatile generalization has driven their broad adoption, yet their deployment is constrained by critical gaps in trustworthiness. First, we conduct a survey to present a comprehensive analysis of what constitute trustworthy foundation model and examined approaches for each requirement from a lifecycle perspective. We identify out-of-distribution (OOD) robustness as the primary technical bottleneck, where performance collapse under distribution shifts invalidates other pillars like fairness and accountability. The nature of this vulnerability is two-fold: first, a reliance on spurious background cues during inference, and second, dimensional collapse in the representation space during fine-tuning. This research aims to mitigate these failures by developing mechanistic insights and algorithmic solutions during foundation model adaptation and deployment phases of the lifecycle.
Keywords:
Machine Learning: Robustness
Machine Learning: Representation learning
Machine Learning: Trustworthy machine learning
AI Ethics, Trust, Fairnes: Trustworthy AI