Wavesplit: End-to-End Speech Separation by Speaker Clustering

Zeghidour, Neil; Grangier, David

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2002

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Wavesplit: End-to-End Speech Separation by Speaker Clustering

Authors: Neil Zeghidour, David Grangier

(Submitted on 20 Feb 2020 (v1), last revised 2 Jul 2020 (this version, v2))

Abstract: We introduce Wavesplit, an end-to-end source separation system. From a single mixture, the model infers a representation for each source and then estimates each source signal given the inferred representations. The model is trained to jointly perform both tasks from the raw waveform. Wavesplit infers a set of source representations via clustering, which addresses the fundamental permutation problem of separation. For speech separation, our sequence-wide speaker representations provide a more robust separation of long, challenging recordings compared to prior work. Wavesplit redefines the state-of-the-art on clean mixtures of 2 or 3 speakers (WSJ0-2/3mix), as well as in noisy and reverberated settings (WHAM/WHAMR). We also set a new benchmark on the recent LibriMix dataset. Finally, we show that Wavesplit is also applicable to other domains, by separating fetal and maternal heart rates from a single abdominal electrocardiogram.

Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD); Machine Learning (stat.ML)
Cite as:	arXiv:2002.08933 [eess.AS]
	(or arXiv:2002.08933v2 [eess.AS] for this version)

Submission history

From: Neil Zeghidour [view email]
[v1] Thu, 20 Feb 2020 18:30:36 GMT (173kb,D)
[v2] Thu, 2 Jul 2020 13:57:33 GMT (257kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2002.08933

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Wavesplit: End-to-End Speech Separation by Speaker Clustering

Submission history