Interpretable Representation Learning for Speech and Audio Signals Based on Relevance Weighting

Agrawal, Purvi; Ganapathy, Sriram

doi:10.1109/TASLP.2020.3030489

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2011

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Interpretable Representation Learning for Speech and Audio Signals Based on Relevance Weighting

Authors: Purvi Agrawal, Sriram Ganapathy

(Submitted on 29 Oct 2020)

Abstract: The learning of interpretable representations from raw data presents significant challenges for time series data like speech. In this work, we propose a relevance weighting scheme that allows the interpretation of the speech representations during the forward propagation of the model itself. The relevance weighting is achieved using a sub-network approach that performs the task of feature selection. A relevance sub-network, applied on the output of first layer of a convolutional neural network model operating on raw speech signals, acts as an acoustic filterbank (FB) layer with relevance weighting. A similar relevance sub-network applied on the second convolutional layer performs modulation filterbank learning with relevance weighting. The full acoustic model consisting of relevance sub-networks, convolutional layers and feed-forward layers is trained for a speech recognition task on noisy and reverberant speech in the Aurora-4, CHiME-3 and VOiCES datasets. The proposed representation learning framework is also applied for the task of sound classification in the UrbanSound8K dataset. A detailed analysis of the relevance weights learned by the model reveals that the relevance weights capture information regarding the underlying speech/audio content. In addition, speech recognition and sound classification experiments reveal that the incorporation of relevance weighting in the neural network architecture improves the performance significantly.

Comments:	arXiv admin note: text overlap with arXiv:2011.00721
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Journal reference:	IEEE Transactions and Audio, Speech and Language Processing, Vol. 28, pp. 2823 - 2836, 2020
DOI:	10.1109/TASLP.2020.3030489
Cite as:	arXiv:2011.02136 [eess.AS]
	(or arXiv:2011.02136v1 [eess.AS] for this version)

Submission history

From: Purvi Agrawal [view email]
[v1] Thu, 29 Oct 2020 20:36:25 GMT (10355kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2011.02136

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Interpretable Representation Learning for Speech and Audio Signals Based on Relevance Weighting

Submission history