VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

Shrivastava, Ayush; Gopalakrishnan, Karthik; Liu, Yang; Piramuthu, Robinson; Tür, Gokhan; Parikh, Devi; Hakkani-Tür, Dilek

Full-text links:

Download:

Current browse context:

cs.CV

< prev | next >

new | recent | 2105

Computer Science > Computer Vision and Pattern Recognition

Title: VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

Authors: Ayush Shrivastava, Karthik Gopalakrishnan, Yang Liu, Robinson Piramuthu, Gokhan Tür, Devi Parikh, Dilek Hakkani-Tür

(Submitted on 25 May 2021 (v1), last revised 16 Mar 2022 (this version, v2))

Abstract: Interactive robots navigating photo-realistic environments need to be trained to effectively leverage and handle the dynamic nature of dialogue in addition to the challenges underlying vision-and-language navigation (VLN). In this paper, we present VISITRON, a multi-modal Transformer-based navigator better suited to the interactive regime inherent to Cooperative Vision-and-Dialog Navigation (CVDN). VISITRON is trained to: i) identify and associate object-level concepts and semantics between the environment and dialogue history, ii) identify when to interact vs. navigate via imitation learning of a binary classification head. We perform extensive pre-training and fine-tuning ablations with VISITRON to gain empirical insights and improve performance on CVDN. VISITRON's ability to identify when to interact leads to a natural generalization of the game-play mode introduced by Roman et al. (arXiv:2005.00728) for enabling the use of such models in different environments. VISITRON is competitive with models on the static CVDN leaderboard and attains state-of-the-art performance on the Success weighted by Path Length (SPL) metric.

Comments:	Accepted at Findings of the Annual Meeting of the Association for Computational Linguistics (ACL) 2022, previous version accepted at Visually Grounded Interaction and Language (ViGIL) Workshop at NAACL 2021
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG); Robotics (cs.RO)
ACM classes:	I.2.9
Cite as:	arXiv:2105.11589 [cs.CV]
	(or arXiv:2105.11589v2 [cs.CV] for this version)

Submission history

From: Karthik Gopalakrishnan [view email]
[v1] Tue, 25 May 2021 00:21:54 GMT (1239kb,D)
[v2] Wed, 16 Mar 2022 03:03:00 GMT (14044kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2105.11589

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Computer Vision and Pattern Recognition

Title: VISITRON: Visual Semantics-Aligned Interactively Trained Object-Navigator

Submission history