Explicit Prosodic Modelling and Deep Speaker Embedding Learning for Non-standard Voice Conversion

Wang, Disong; Liu, Songxiang; Sun, Lifa; Wu, Xixin; Liu, Xunying; Meng, Helen

Full-text links:

Download:

Current browse context:

eess.AS

< prev | next >

new | recent | 2011

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Explicit Prosodic Modelling and Deep Speaker Embedding Learning for Non-standard Voice Conversion

Authors: Disong Wang, Songxiang Liu, Lifa Sun, Xixin Wu, Xunying Liu, Helen Meng

(Submitted on 3 Nov 2020 (this version), latest version 17 Jun 2021 (v2))

Abstract: Though significant progress has been made for the voice conversion (VC) of standard speech, VC for non-standard speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for atypical prosody while maintaining speaker identity. To address this issue, we propose a VC system with explicit prosody modelling and deep speaker embedding (DSE) learning. First, a speech-encoder strives to extract robust phoneme embeddings from non-standard speech. Second, a prosody corrector takes in phoneme embeddings to infer standard phoneme duration and pitch values. Third, a conversion model takes phoneme embeddings and standard prosody features as inputs to generate the converted speech, conditioned on the target DSE that is learned via speaker encoder or speaker adaptation. Extensive experiments demonstrate that speaker encoder based conversion model can significantly reduce dysarthric and non-native pronunciation patterns to generate near-normal and near-native speech respectively, and speaker adaptation can achieve higher speaker similarity.

Comments:	Submitted to ICASSP2021
Subjects:	Audio and Speech Processing (eess.AS); Computation and Language (cs.CL); Sound (cs.SD)
Cite as:	arXiv:2011.01678 [eess.AS]
	(or arXiv:2011.01678v1 [eess.AS] for this version)

Submission history

From: Disong Wang [view email]
[v1] Tue, 3 Nov 2020 13:08:53 GMT (876kb,D)
[v2] Thu, 17 Jun 2021 12:50:49 GMT (835kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2011.01678v1

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: Explicit Prosodic Modelling and Deep Speaker Embedding Learning for Non-standard Voice Conversion

Submission history