Perturbation-Resilient Autonomous Navigation with Distributionally Robust Reinforcement Learning

Perturbation-Resilient Autonomous Navigation with Distributionally Robust Reinforcement Learning

Zhaofan Zhang, Minghao Yang, Sihong Xie, Hui Xiong

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
AI and Robotics. Pages 7619-7627. https://doi.org/10.24963/ijcai.2026/847

The robustness of autonomous vehicles such as drones and Unmanned Surface Vehicles (USV) is crucial when facing unknown and complex marine environments, especially when heteroscedastic observational noise poses significant challenges to sensor-based navigation tasks. Recently, Distributional Reinforcement Learning (DistRL) has shown promising results in some challenging autonomous navigation tasks without prior environmental information. However, these methods overlook situations where noise patterns vary across different environmental conditions, hindering safe navigation and disrupting the learning of value functions. To address the problem, we propose DRIQN to integrate Distributionally Robust Optimization (DRO) with implicit quantile networks to optimize worst-case performance under natural environmental conditions. Leveraging explicit subgroup modeling in the replay buffer, DRIQN incorporates heterogeneous noise sources and target robustness-critical scenarios. Experimental results based on the risk-sensitive environment demonstrate that DRIQN significantly outperforms state-of-the-art meth- ods, achieving +13.51% success rate, -12.28% collision rate and +35.46% for time saving, +27.99% for energy saving, compared with the runner-up.
Keywords:
AIR: Robot control, planning, and execution with guarantees
AIR: Generative AI, robotic foundation models, and reinforcement learning
Robot control, planning, and execution with guarantees: Safe and robust control under uncertainty
Safety, trustworthiness, generalizability, and evaluation: Theoretical and algorithmic guarantees on safety, robustness, and out-of-distribution generalization