Multi-task Language Modeling for Improving Speech Recognition of Rare Words

Yang, Chao-Han Huck; Liu, Linda; Gandhe, Ankur; Gu, Yile; Raju, Anirudh; Filimonov, Denis; Bulyko, Ivan

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2011

Computer Science > Computation and Language

Title: Multi-task Language Modeling for Improving Speech Recognition of Rare Words

Authors: Chao-Han Huck Yang, Linda Liu, Ankur Gandhe, Yile Gu, Anirudh Raju, Denis Filimonov, Ivan Bulyko

(Submitted on 23 Nov 2020 (v1), revised 25 Nov 2020 (this version, v2), latest version 11 Sep 2021 (v4))

Abstract: End-to-end automatic speech recognition (ASR) systems are increasingly popular due to their relative architectural simplicity and competitive performance. However, even though the average accuracy of these systems may be high, the performance on rare content words often lags behind hybrid ASR systems. To address this problem, second-pass rescoring is often applied. In this paper, we propose a second-pass system with multi-task learning, utilizing semantic targets (such as intent and slot prediction) to improve speech recognition performance. We show that our rescoring model with trained with these additional tasks outperforms the baseline rescoring model, trained with only the language modeling task, by 1.4% on a general test and by 2.6% on a rare word test set in term of word-error-rate relative (WERR).

Comments:	Submitted to ICASSP 2021
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Neural and Evolutionary Computing (cs.NE); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2011.11715 [cs.CL]
	(or arXiv:2011.11715v2 [cs.CL] for this version)

Submission history

From: C.-H. Huck Yang [view email]
[v1] Mon, 23 Nov 2020 20:40:44 GMT (96kb,D)
[v2] Wed, 25 Nov 2020 03:12:54 GMT (96kb,D)
[v3] Fri, 2 Apr 2021 20:31:00 GMT (100kb,D)
[v4] Sat, 11 Sep 2021 21:58:38 GMT (101kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2011.11715v2

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Multi-task Language Modeling for Improving Speech Recognition of Rare Words

Submission history