MoTRa: Motion-Aware Target Representation Learning for End-to-End Multi-Object Tracking
MoTRa: Motion-Aware Target Representation Learning for End-to-End Multi-Object Tracking
Yuanzhou Huang, Songwei Pei, Shuhuai Wang, Bingfeng Liu, Qian Li, Shangguang Wang
Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 1224-1232.
https://doi.org/10.24963/ijcai.2026/137
Multi-object tracking (MOT) has long faced challenges with identity switches, especially for targets with low appearance discriminability and complex motion. Existing end-to-end trackers typically enhance robustness by modeling long-term temporal information across target-level representations, yet these representations remain insufficiently discriminative for spatially and visually similar targets. In this paper, we present MoTRa, a Motion-aware Target Representation learning framework for end-to-end MOT that enriches target-level representations with adaptive motion cues. As a result, each representation adaptively balances motion cues and appearance-related content cues under varying conditions, enhancing its discriminability for tracking. To achieve this, we propose a feature fusion and alignment module that extracts motion cues and adaptively fuses them with content features into target representations. To further regularize the dynamically fused representations, we propose a target-specific contrastive learning strategy that promotes intra-trajectory consistency and inter-target separability. Experimental results demonstrate the effectiveness of our method, achieving competitive performance on multiple benchmarks.
Keywords:
Computer Vision: Motion and tracking
Computer Vision: Representation learning
Computer Vision: Video analysis and understanding
