Transferable Attacks on Open-Vocabulary Video Instance Segmentation via Dual-Objective Triggers

Transferable Attacks on Open-Vocabulary Video Instance Segmentation via Dual-Objective Triggers

Minghao Shou, Kesen Wang, Tong Zhang, Han Bao, Zonghui Wang

Proceedings of the Thirty-Fifth International Joint Conference on Artificial Intelligence
Main Track. Pages 1613-1621. https://doi.org/10.24963/ijcai.2026/180

Open‑vocabulary video instance segmentation (OV‑VIS) couples spatial‑temporal reasoning with language grounding, yet its adversarial robustness has remained unexplored. We present the Dual-Objective Triggers (DOT), the first transferable attack on OV-VIS that simultaneously exploits the vision–language coupling and temporal coherence. DOT deploys a Dual Semantic Perturbation Module that overlays two complementary triggers: a Semantic Suppression Trigger erases the alignment between the true object and the query, while a Plausible Replacement Trigger steers the tracker toward a phantom trajectory that is visually plausible and text‑consistent. To amplify cross‑model transferability without sacrificing perceptual fidelity, we introduce Phase‑Guided Adversarial Training, which injects perturbations primarily in the phase spectrum while blending amplitudes with clean references. Extensive experiments on four state‑of‑the‑art OV‑VIS implementations demonstrate that DOT reduces mAP by up to 69.3% and raises attack success rate by up to 98%, outperforming the strongest baselines by a factor of 1.6× on average, while maintaining a PSNR of 52.49 dB, thus exposing critical security vulnerabilities and laying a foundation for future research on robust and trustworthy vision–language systems.
Keywords:
Computer Vision: Adversarial learning, adversarial attack and defense methods
Computer Vision: Video analysis and understanding
Computer Vision: Vision, language and reasoning
Machine Learning: Adversarial machine learning
Machine Learning: Multi-modal learning