Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English with Transfer Learning

Shibano, Toshiko; Zhang, Xinyi; Li, Mia Taige; Cho, Haejin; Sullivan, Peter; Abdul-Mageed, Muhammad

Full-text links:

Download:

Current browse context:

eess

< prev | next >

new | recent | 2110

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English with Transfer Learning

Authors: Toshiko Shibano (1), Xinyi Zhang (1), Mia Taige Li (1), Haejin Cho (1), Peter Sullivan (1), Muhammad Abdul-Mageed (1) ((1) University of British Columbia)

(Submitted on 1 Oct 2021 (v1), last revised 15 Oct 2021 (this version, v3))

Abstract: To address the performance gap of English ASR models on L2 English speakers, we evaluate fine-tuning of pretrained wav2vec 2.0 models (Baevski et al., 2020; Xu et al., 2021) on L2-ARCTIC, a non-native English speech corpus (Zhao et al., 2018) under different training settings. We compare \textbf{(a)} models trained with a combination of diverse accents to ones trained with only specific accents and \textbf{(b)} results from different single-accent models. Our experiments demonstrate the promise of developing ASR models for non-native English speakers, even with small amounts of L2 training data and even without a language model. Our models also excel in the zero-shot setting where we train on multiple L2 datasets and test on a blind L2 test set.

Comments:	All authors contributed equally. Paper accepted to International Conference on Natural Language and Speech Processing 2021 (ICNLSP 2021)
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2110.00678 [eess.AS]
	(or arXiv:2110.00678v3 [eess.AS] for this version)

Submission history

From: Peter Sullivan [view email]
[v1] Fri, 1 Oct 2021 23:11:00 GMT (5462kb,D)
[v2] Wed, 13 Oct 2021 19:45:09 GMT (5462kb,D)
[v3] Fri, 15 Oct 2021 02:43:42 GMT (5462kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2110.00678

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Speech Technology for Everyone: Automatic Speech Recognition for Non-Native English with Transfer Learning

Submission history