ViSA-Gait: Leveraging Vision Foundation Models for Semantic Anchored Gait Recognition
ViSA-Gait: Leveraging Vision Foundation Models for Semantic Anchored Gait Recognition
Xiangru Li, Deqiang Yin, Yifan Xie, Guojian Li, Zebang Cheng, Fei Ma
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 1351-1359.
https://doi.org/10.24963/ijcai.2026/151
Gait recognition has achieved remarkable success in constrained environments, yet its performance often degrades significantly in cross-domain and cross-vertical-view scenarios. This is primarily due to the fact that domain-specific silhouette geometry causes models to overfit to extrinsic geometric characteristics rather than learning generalizable motion patterns. To address this issue, we present ViSA-Gait, which utilizes Vision Foundation Models (VFMs) as a “Semantic Compass” for stable universal human body knowledge guidance, enabling a lightweight backbone to robustly capture fine-grained gait dynamics. Specifically, we introduce a token-based distillation mechanism where a Spatial Token Learner (STL) and a Temporal Token Learner (TTL) filter dense VFM features into motion-consistent descriptors. These descriptors are then adaptively injected into a lightweight 3D-CNN backbone via a Gated Cross-Attention (GCA) mechanism, functioning as “Semantic Anchors” that regularize the feature space and guide the model to focus on intrinsic body motion. Extensive experiments demonstrate that ViSA-Gait achieves SOTA performance on cross-domain and cross-vertical-view benchmarks while remaining competitive within-domain, offering a new perspective on bridging the gap between geometry-based analysis and universal semantic understanding.
Keywords:
Computer Vision: Low-level Vision
Computer Vision: Machine learning for vision
Computer Vision: Recognition (object detection, categorization)
Computer Vision: Representation learning
