Speaker Recognition Based on Deep Learning: An Overview

Bai, Zhongxin; Zhang, Xiao-Lei

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2012

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Speaker Recognition Based on Deep Learning: An Overview

Authors: Zhongxin Bai, Xiao-Lei Zhang

(Submitted on 2 Dec 2020 (v1), last revised 4 Apr 2021 (this version, v2))

Abstract: Speaker recognition is a task of identifying persons from their voices. Recently, deep learning has dramatically revolutionized speaker recognition. However, there is lack of comprehensive reviews on the exciting progress.
In this paper, we review several major subtasks of speaker recognition, including speaker verification, identification, diarization, and robust speaker recognition, with a focus on deep-learning-based methods. Because the major advantage of deep learning over conventional methods is its representation ability, which is able to produce highly abstract embedding features from utterances, we first pay close attention to deep-learning-based speaker feature extraction, including the inputs, network structures, temporal pooling strategies, and objective functions respectively, which are the fundamental components of many speaker recognition subtasks. Then, we make an overview of speaker diarization, with an emphasis of recent supervised, end-to-end, and online diarization. Finally, we survey robust speaker recognition from the perspectives of domain adaptation and speech enhancement, which are two major approaches of dealing with domain mismatch and noise problems. Popular and recently released corpora are listed at the end of the paper.

Subjects:	Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2012.00931 [eess.AS]
	(or arXiv:2012.00931v2 [eess.AS] for this version)

Submission history

From: Zhongxin Bai [view email]
[v1] Wed, 2 Dec 2020 02:24:45 GMT (17107kb,D)
[v2] Sun, 4 Apr 2021 01:59:30 GMT (3668kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2012.00931

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Speaker Recognition Based on Deep Learning: An Overview

Submission history