Learning towards Selective Data Augmentation for Dialogue Generation

Chen, Xiuying; Li, Mingzhe; Zhang, Jiayi; Xia, Xiaoqiang; Wei, Chen; Cui, Jianwei; Gao, Xin; Zhang, Xiangliang; Yan, Rui

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 2303

Change to browse by:

Computer Science > Computation and Language

Title: Learning towards Selective Data Augmentation for Dialogue Generation

Authors: Xiuying Chen, Mingzhe Li, Jiayi Zhang, Xiaoqiang Xia, Chen Wei, Jianwei Cui, Xin Gao, Xiangliang Zhang, Rui Yan

(Submitted on 17 Mar 2023)

Abstract: As it is cumbersome and expensive to acquire a huge amount of data for training neural dialog models, data augmentation is proposed to effectively utilize existing training samples. However, current data augmentation techniques on the dialog generation task mostly augment all cases in the training dataset without considering the intrinsic attributes between different cases. We argue that not all cases are beneficial for augmentation task, and the cases suitable for augmentation should obey the following two attributes: (1) low-quality (the dialog model cannot generate a high-quality response for the case), (2) representative (the case should represent the property of the whole dataset). Herein, we explore this idea by proposing a Selective Data Augmentation framework (SDA) for the response generation task. SDA employs a dual adversarial network to select the lowest quality and most representative data points for augmentation in one stage. Extensive experiments conducted on two publicly available datasets, i.e., DailyDialog and OpenSubtitles, show that our framework can improve the response generation performance with respect to various metrics.

Comments:	9 pages, 4 figures, AAAI 2023
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2303.09719 [cs.CL]
	(or arXiv:2303.09719v1 [cs.CL] for this version)

Submission history

From: Xiuying Chen [view email]
[v1] Fri, 17 Mar 2023 01:26:39 GMT (5269kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2303.09719

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computation and Language

Title: Learning towards Selective Data Augmentation for Dialogue Generation

Submission history