A Post Auto-regressive GAN Vocoder Focused on Spectrum Fracture

Lu, Zhenxing; He, Mengnan; Zhang, Ruixiong; Gong, Caixia

Full-text links:

Download:

Source

Current browse context:

eess.AS

< prev | next >

new | recent | 2204

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: A Post Auto-regressive GAN Vocoder Focused on Spectrum Fracture

Authors: Zhenxing Lu, Mengnan He, Ruixiong Zhang, Caixia Gong

(Submitted on 12 Apr 2022 (v1), last revised 16 Feb 2023 (this version, v2))

Abstract: Generative adversarial networks (GANs) have been indicated their superiority in usage of the real-time speech synthesis. Nevertheless, most of them make use of deep convolutional layers as their backbone, which may cause the absence of previous signal information. However, the generation of speech signals invariably require preceding waveform samples in its reconstruction, as the lack of this can lead to artifacts in generated speech. To address this conflict, in this paper, we propose an improved model: a post auto-regressive (AR) GAN vocoder with a self-attention layer, which merging self-attention in an AR loop. It will not participate in inference, but can assist the generator to learn temporal dependencies within frames in training. Furthermore, an ablation study was done to confirm the contribution of each part. Systematic experiments show that our model leads to a consistent improvement on both objective and subjective evaluation performance.

Comments:	Experimental parts should be improved
Subjects:	Audio and Speech Processing (eess.AS); Sound (cs.SD)
Cite as:	arXiv:2204.06086 [eess.AS]
	(or arXiv:2204.06086v2 [eess.AS] for this version)

Submission history

From: Zhenxing Lu [view email]
[v1] Tue, 12 Apr 2022 21:33:10 GMT (1439kb,D)
[v2] Thu, 16 Feb 2023 05:11:10 GMT (0kb,I)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> eess > arXiv:2204.06086

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Electrical Engineering and Systems Science > Audio and Speech Processing

Title: A Post Auto-regressive GAN Vocoder Focused on Spectrum Fracture

Submission history