Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents

Dubey, Priyank; Shah, Bilal

Full-text links:

Download:

PDF only

Current browse context:

cs.CL

< prev | next >

new | recent | 2204

Computer Science > Computation and Language

Title: Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents

Authors: Priyank Dubey, Bilal Shah

(Submitted on 3 Apr 2022)

Abstract: Automated Speech Recognition (ASR) is an interdisciplinary application of computer science and linguistics that enable us to derive the transcription from the uttered speech waveform. It finds several applications in Military like High-performance fighter aircraft, helicopters, air-traffic controller. Other than military speech recognition is used in healthcare, persons with disabilities and many more. ASR has been an active research area. Several models and algorithms for speech to text (STT) have been proposed. One of the most recent is Mozilla Deep Speech, it is based on the Deep Speech research paper by Baidu. Deep Speech is a state-of-art speech recognition system is developed using end-to-end deep learning, it is trained using well-optimized Recurrent Neural Network (RNN) training system utilizing multiple Graphical Processing Units (GPUs). This training is mostly done using American-English accent datasets, which results in poor generalizability to other English accents. India is a land of vast diversity. This can even be seen in the speech, there are several English accents which vary from state to state. In this work, we have used transfer learning approach using most recent Deep Speech model i.e., deepspeech-0.9.3 to develop an end-to-end speech recognition system for Indian-English accents. This work utilizes fine-tuning and data argumentation to further optimize and improve the Deep Speech ASR system. Indic TTS data of Indian-English accents is used for transfer learning and fine-tuning the pre-trained Deep Speech model. A general comparison is made among the untrained model, our trained model and other available speech recognition services for Indian-English Accents.

Subjects:	Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2204.00977 [cs.CL]
	(or arXiv:2204.00977v1 [cs.CL] for this version)

Submission history

From: Priyank Dubey Mr. [view email]
[v1] Sun, 3 Apr 2022 03:11:21 GMT (281kb)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2204.00977

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Deep Speech Based End-to-End Automated Speech Recognition (ASR) for Indian-English Accents

Submission history