Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

Jie Wu; Wei Zhang; Guanbin Li; Wenhao Wu; Xiao Tan; Yingying Li; Errui Ding; Liang Lin

doi:10.24963/ijcai.2021/162

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

Jie Wu, Wei Zhang, Guanbin Li, Wenhao Wu, Xiao Tan, Yingying Li, Errui Ding, Liang Lin

Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence

Main Track. Pages 1172-1178. https://doi.org/10.24963/ijcai.2021/162

PDF BibTeX

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i.e., a sequence of bounding boxes at consecutive times) that encloses the abnormal event, with only coarse video-level annotations as supervision during training. To address this challenging task, we propose a dual-branch network which takes as input the proposals with multi-granularities in both spatial-temporal domains. Each branch employs a relationship reasoning module to capture the correlation between tubes/videolets, which can provide rich contextual information and complex entity relationships for the concept learning of abnormal behaviors. Mutually-guided Progressive Refinement framework is set up to employ dual-path mutual guidance in a recurrent manner, iteratively sharing auxiliary supervision information across branches. It impels the learned concepts of each branch to serve as a guide for its counterpart, which progressively refines the corresponding branch and the whole framework. Furthermore, we contribute two datasets, i.e., ST-UCF-Crime and STRA, consisting of videos containing spatio-temporal abnormal annotations to serve as the benchmarks for WSSTAD. We conduct extensive qualitative and quantitative evaluations to demonstrate the effectiveness of the proposed approach and analyze the key factors that contribute more to handle this task.

Keywords:

Computer Vision: Video: Events, Activities and Surveillance

Machine Learning: Weakly Supervised Learning