Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation

Zmolikova, Katerina; Delcroix, Marc; Burget, Lukáš; Nakatani, Tomohiro; Černocký, Jan "Honza"

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2011

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation

Authors: Katerina Zmolikova, Marc Delcroix, Lukáš Burget, Tomohiro Nakatani, Jan "Honza" Černocký

(Submitted on 24 Nov 2020)

Abstract: In this paper, we propose a method combining variational autoencoder model of speech with a spatial clustering approach for multi-channel speech separation. The advantage of integrating spatial clustering with a spectral model was shown in several works. As the spectral model, previous works used either factorial generative models of the mixed speech or discriminative neural networks. In our work, we combine the strengths of both approaches, by building a factorial model based on a generative neural network, a variational autoencoder. By doing so, we can exploit the modeling power of neural networks, but at the same time, keep a structured model. Such a model can be advantageous when adapting to new noise conditions as only the noise part of the model needs to be modified. We show experimentally, that our model significantly outperforms previous factorial model based on Gaussian mixture model (DOLPHIN), performs comparably to integration of permutation invariant training with spatial clustering, and enables us to easily adapt to new noise conditions. The code for the method is available at this https URL

Comments:	8 pages, 3 figures, to be published in SLT2021
Subjects:	Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2011.11984 [eess.AS]
	(or arXiv:2011.11984v1 [eess.AS] for this version)

Submission history

From: Katerina Zmolikova [view email]
[v1] Tue, 24 Nov 2020 09:28:46 GMT (311kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2011.11984

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Integration of variational autoencoder and spatial clustering for adaptive multi-channel neural speech separation

Submission history