Learning Sound Events from Webly Labeled Data

Learning Sound Events from Webly Labeled Data

Anurag Kumar, Ankit Shah, Alexander Hauptmann, Bhiksha Raj

Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence
Main track. Pages 2772-2778. https://doi.org/10.24963/ijcai.2019/384

In the last couple of years, weakly labeled learning has turned out to be an exciting approach for audio event detection. In this work, we introduce webly labeled learning for sound events which aims to remove human supervision altogether from the learning process. We first develop a method of obtaining labeled audio data from the web (albeit noisy), in which no manual labeling is involved. We then describe methods to efficiently learn from these webly labeled audio recordings. In our proposed system, WeblyNet, two deep neural networks co-teach each other to robustly learn from webly labeled data, leading to around 17% relative improvement over the baseline method. The method also involves transfer learning to obtain efficient representations.
Keywords:
Machine Learning: Classification
Machine Learning: Multi-instance;Multi-label;Multi-view learning
Machine Learning: Deep Learning
Machine Learning Applications: Applications of Unsupervised Learning