Structure-Invariant Testing for Machine Translation

He, Pinjia; Meister, Clara; Su, Zhendong

Full-text links:

Download:

Current browse context:

cs.SE

< prev | next >

new | recent | 1907

Computer Science > Software Engineering

Title: Structure-Invariant Testing for Machine Translation

Authors: Pinjia He, Clara Meister, Zhendong Su

(Submitted on 19 Jul 2019 (v1), last revised 14 Jul 2020 (this version, v3))

Abstract: In recent years, machine translation software has increasingly been integrated into our daily lives. People routinely use machine translation for various applications, such as describing symptoms to a foreign doctor and reading political news in a foreign language. However, the complexity and intractability of neural machine translation (NMT) models that power modern machine translation make the robustness of these systems difficult to even assess, much less guarantee. Machine translation systems can return inferior results that lead to misunderstanding, medical misdiagnoses, threats to personal safety, or political conflicts. Despite its apparent importance, validating the robustness of machine translation systems is very difficult and has, therefore, been much under-explored.
To tackle this challenge, we introduce structure-invariant testing (SIT), a novel metamorphic testing approach for validating machine translation software. Our key insight is that the translation results of "similar" source sentences should typically exhibit similar sentence structures. Specifically, SIT (1) generates similar source sentences by substituting one word in a given sentence with semantically similar, syntactically equivalent words; (2) represents sentence structure by syntax parse trees (obtained via constituency or dependency parsing); (3) reports sentence pairs whose structures differ quantitatively by more than some threshold. To evaluate SIT, we use it to test Google Translate and Bing Microsoft Translator with 200 source sentences as input, which led to 64 and 70 buggy issues with 69.5\% and 70\% top-1 accuracy, respectively. The translation errors are diverse, including under-translation, over-translation, incorrect modification, word/phrase mistranslation, and unclear logic.

Comments:	Accepted at ICSE 2020
Subjects:	Software Engineering (cs.SE); Computation and Language (cs.CL)
Cite as:	arXiv:1907.08710 [cs.SE]
	(or arXiv:1907.08710v3 [cs.SE] for this version)

Submission history

From: Pinjia He [view email]
[v1] Fri, 19 Jul 2019 22:20:01 GMT (3382kb,D)
[v2] Sat, 24 Aug 2019 17:37:36 GMT (6124kb,D)
[v3] Tue, 14 Jul 2020 10:30:36 GMT (1581kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1907.08710

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Software Engineering

Title: Structure-Invariant Testing for Machine Translation

Submission history