FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments
FILD-Nav:Vision-and-Language Navigation with Instruction Landmark Features in Continuous Environments
Chuangye Hu, Lulu Liu, Huaiwei Si, Yawen Zhao, Nan Ding
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
AI and Robotics. Pages 7566-7574.
https://doi.org/10.24963/ijcai.2026/841
Vision-and-language navigation (VLN) requires agents to follow natural language instructions to navigate autonomously in continuous environments. However, existing approaches often lack high-level semantic guidance in waypoint prediction and explicit language–landmark alignment in cross-modal planning. To address these limitations, we propose FILD-Nav, a vision-and-language navigation framework that integrates instruction landmark features. FILD-Nav extracts task-relevant landmarks from instructions and incorporates landmark semantics into both waypoint prediction and topological planning. Specifically, landmark-guided waypoint prediction improves waypoint relevance, while landmark-enhanced cross-modal planning enables more effective long-horizon navigation. Extensive experiments on the VLN-CE benchmark demonstrate that FILD-Nav consistently outperforms prior methods, achieving improvements of 2% in Success Rate (SR), 3% in Success weighted by Path Length (SPL), and 7% in Oracle Success Rate (OSR), particularly in unseen environments.
Keywords:
Robot control, planning, and execution with guarantees: Architectures connecting high-level intent and constraints to low-level trajectories
Learning to understand, generalize, and explain actions: Robust generalization and transfer across tasks, objects, environments, embodiments, and long-horizon scenarios
AIR: Robot control, planning, and execution with guarantees
Foundations of human–robot interaction and assistance: Learning and inference methods for aligning robot behavior with human instructions, demonstrations, and feedback
