Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss

Sato, Hiroshi; Masumura, Ryo; Ochiai, Tsubasa; Delcroix, Marc; Moriya, Takafumi; Ashihara, Takanori; Shinayama, Kentaro; Mizuno, Saki; Ihori, Mana; Tanaka, Tomohiro; Hojo, Nobukatsu

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2305

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss

Authors: Hiroshi Sato, Ryo Masumura, Tsubasa Ochiai, Marc Delcroix, Takafumi Moriya, Takanori Ashihara, Kentaro Shinayama, Saki Mizuno, Mana Ihori, Tomohiro Tanaka, Nobukatsu Hojo

(Submitted on 24 May 2023)

Abstract: Self-supervised learning (SSL) is the latest breakthrough in speech processing, especially for label-scarce downstream tasks by leveraging massive unlabeled audio data. The noise robustness of the SSL is one of the important challenges to expanding its application. We can use speech enhancement (SE) to tackle this issue. However, the mismatch between the SE model and SSL models potentially limits its effect. In this work, we propose a new SE training criterion that minimizes the distance between clean and enhanced signals in the feature representation of the SSL model to alleviate the mismatch. We expect that the loss in the SSL domain could guide SE training to preserve or enhance various levels of characteristics of the speech signals that may be required for high-level downstream tasks. Experiments show that our proposal improves the performance of an SE and SSL pipeline on five downstream tasks with noisy input while maintaining the SE performance.

Comments:	4 pages , 2 figures, Accepted to Interspeech 2023
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2305.14723 [eess.AS]
	(or arXiv:2305.14723v1 [eess.AS] for this version)

Submission history

From: Hiroshi Sato [view email]
[v1] Wed, 24 May 2023 05:00:30 GMT (1184kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2305.14723

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Downstream Task Agnostic Speech Enhancement with Self-Supervised Representation Loss

Submission history