Some voices are too common: Building fair speech recognition systems using the Common Voice dataset

Maison, Lucas; Estève, Yannick

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2306

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Some voices are too common: Building fair speech recognition systems using the Common Voice dataset

Authors: Lucas Maison, Yannick Estève

(Submitted on 1 Jun 2023)

Abstract: Automatic speech recognition (ASR) systems become increasingly efficient thanks to new advances in neural network training like self-supervised learning. However, they are known to be unfair toward certain groups, for instance, people speaking with an accent. In this work, we use the French Common Voice dataset to quantify the biases of a pre-trained wav2vec~2.0 model toward several demographic groups. By fine-tuning the pre-trained model on a variety of fixed-size, carefully crafted training sets, we demonstrate the importance of speaker diversity. We also run an in-depth analysis of the Common Voice corpus and identify important shortcomings that should be taken into account by users of this dataset.

Comments:	5 pages, 3 figures. Accepted to Interspeech 2023
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Machine Learning (cs.LG); Sound (cs.SD)
Cite as:	arXiv:2306.03773 [eess.AS]
	(or arXiv:2306.03773v1 [eess.AS] for this version)

Submission history

From: Lucas Maison [view email]
[v1] Thu, 1 Jun 2023 11:42:34 GMT (309kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2306.03773

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Some voices are too common: Building fair speech recognition systems using the Common Voice dataset

Submission history