Gaussian-smoothed Imbalance Data Improves Speech Emotion Recognition

Liang, Xuefeng; Jiang, Hexin; Xu, Wenxin; Zhou, Ying

Full-text links:

Download:

Current browse context:

cs.SD

< prev | next >

new | recent | 2302

Computer Science > Sound

Title: Gaussian-smoothed Imbalance Data Improves Speech Emotion Recognition

Authors: Xuefeng Liang, Hexin Jiang, Wenxin Xu, Ying Zhou

(Submitted on 17 Feb 2023)

Abstract: In speech emotion recognition tasks, models learn emotional representations from datasets. We find the data distribution in the IEMOCAP dataset is very imbalanced, which may harm models to learn a better representation. To address this issue, we propose a novel Pairwise-emotion Data Distribution Smoothing (PDDS) method. PDDS considers that the distribution of emotional data should be smooth in reality, then applies Gaussian smoothing to emotion-pairs for constructing a new training set with a smoother distribution. The required new data are complemented using the mixup augmentation. As PDDS is model and modality agnostic, it is evaluated with three SOTA models on the IEMOCAP dataset. The experimental results show that these models are improved by 0.2\% - 4.8\% and 1.5\% - 5.9\% in terms of WA and UA. In addition, an ablation study demonstrates that the key advantage of PDDS is the reasonable data distribution rather than a simple data augmentation.

Comments:	5 pages
Subjects:	Sound (cs.SD); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2302.08650 [cs.SD]
	(or arXiv:2302.08650v1 [cs.SD] for this version)

Submission history

From: Xuefeng Liang [view email]
[v1] Fri, 17 Feb 2023 01:50:46 GMT (3117kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2302.08650

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Sound

Title: Gaussian-smoothed Imbalance Data Improves Speech Emotion Recognition

Submission history