Multimodal Attention Networks for Low-Level Vision-and-Language Navigation

Landi, Federico; Baraldi, Lorenzo; Cornia, Marcella; Corsini, Massimiliano; Cucchiara, Rita

Full-text links:

Download:

Current browse context:

cs.CL

< prev | next >

new | recent | 1911

Computer Science > Computer Vision and Pattern Recognition

Title: Multimodal Attention Networks for Low-Level Vision-and-Language Navigation

Authors: Federico Landi, Lorenzo Baraldi, Marcella Cornia, Massimiliano Corsini, Rita Cucchiara

(Submitted on 27 Nov 2019 (v1), last revised 30 Jul 2021 (this version, v3))

Abstract: Vision-and-Language Navigation (VLN) is a challenging task in which an agent needs to follow a language-specified path to reach a target destination. The goal gets even harder as the actions available to the agent get simpler and move towards low-level, atomic interactions with the environment. This setting takes the name of low-level VLN. In this paper, we strive for the creation of an agent able to tackle three key issues: multi-modality, long-term dependencies, and adaptability towards different locomotive settings. To that end, we devise "Perceive, Transform, and Act" (PTA): a fully-attentive VLN architecture that leaves the recurrent approach behind and the first Transformer-like architecture incorporating three different modalities - natural language, images, and low-level actions for the agent control. In particular, we adopt an early fusion strategy to merge lingual and visual information efficiently in our encoder. We then propose to refine the decoding phase with a late fusion extension between the agent's history of actions and the perceptual modalities. We experimentally validate our model on two datasets: PTA achieves promising results in low-level VLN on R2R and achieves good performance in the recently proposed R4R benchmark. Our code is publicly available at this https URL

Comments:	Computer Vision and Image Understanding (CVIU)
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:1911.12377 [cs.CV]
	(or arXiv:1911.12377v3 [cs.CV] for this version)

Submission history

From: Federico Landi [view email]
[v1] Wed, 27 Nov 2019 19:00:24 GMT (747kb,D)
[v2] Mon, 20 Jul 2020 07:30:16 GMT (747kb,D)
[v3] Fri, 30 Jul 2021 09:13:11 GMT (922kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:1911.12377

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: Multimodal Attention Networks for Low-Level Vision-and-Language Navigation

Submission history