Monotonic Chunkwise Attention

Chiu, Chung-Cheng; Raffel, Colin

Full-text links:

Download:

Current browse context:

stat

< prev | next >

new | recent | 1712

Computer Science > Computation and Language

Title: Monotonic Chunkwise Attention

Authors: Chung-Cheng Chiu, Colin Raffel

(Submitted on 14 Dec 2017 (v1), last revised 23 Feb 2018 (this version, v2))

Abstract: Sequence-to-sequence models with soft attention have been successfully applied to a wide variety of problems, but their decoding process incurs a quadratic time and space cost and is inapplicable to real-time sequence transduction. To address these issues, we propose Monotonic Chunkwise Attention (MoChA), which adaptively splits the input sequence into small chunks over which soft attention is computed. We show that models utilizing MoChA can be trained efficiently with standard backpropagation while allowing online and linear-time decoding at test time. When applied to online speech recognition, we obtain state-of-the-art results and match the performance of a model using an offline soft attention mechanism. In document summarization experiments where we do not expect monotonic alignments, we show significantly improved performance compared to a baseline monotonic attention-based model.

Comments:	ICLR camera-ready version
Subjects:	Computation and Language (cs.CL); Machine Learning (stat.ML)
Cite as:	arXiv:1712.05382 [cs.CL]
	(or arXiv:1712.05382v2 [cs.CL] for this version)

Submission history

From: Chung-Cheng Chiu [view email]
[v1] Thu, 14 Dec 2017 18:29:42 GMT (1037kb,D)
[v2] Fri, 23 Feb 2018 01:35:36 GMT (195kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1712.05382

Download:

Current browse context:

Change to browse by:

References & Citations

Bookmark

Computer Science > Computation and Language

Title: Monotonic Chunkwise Attention

Submission history