Multi-level Attention Fusion Network for Audio-visual Event Recognition

Brousmiche, Mathilde; Rouat, Jean; Dupont, Stéphane

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2106

Computer Science > Computer Vision and Pattern Recognition

Title: Multi-level Attention Fusion Network for Audio-visual Event Recognition

Authors: Mathilde Brousmiche, Jean Rouat, Stéphane Dupont

(Submitted on 12 Jun 2021)

Abstract: Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion network (MAFnet), an architecture that can dynamically fuse visual and audio information for event recognition. Inspired by prior studies in neuroscience, we couple both modalities at different levels of visual and audio paths. Furthermore, the network dynamically highlights a modality at a given time window relevant to classify events. Experimental results in AVE (Audio-Visual Event), UCF51, and Kinetics-Sounds datasets show that the approach can effectively improve the accuracy in audio-visual event classification. Code is available at: this https URL

Comments:	Preprint submitted to the Information Fusion journal in August 2020
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM)
Cite as:	arXiv:2106.06736 [cs.CV]
	(or arXiv:2106.06736v1 [cs.CV] for this version)

Submission history

From: Mathilde Brousmiche [view email]
[v1] Sat, 12 Jun 2021 10:24:52 GMT (2167kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2106.06736

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: Multi-level Attention Fusion Network for Audio-visual Event Recognition

Submission history