VLN(sic)BERT: A Recurrent Vision-and-Language BERT for Navigation
Loading...
Date
Authors
Hong, Yicong
Wu, Qi
Qi, Yuankai
Rodriguez Opazo, Cristian
Gould, Stephen
Journal Title
Journal ISSN
Volume Title
Publisher
IEEE Computer Society
Abstract
Accuracy of many visiolinguistic tasks has benefited significantly from the application of vision-and-language (V&L) BERT. However, its application for the task of vision-and-language navigation (VLN) remains limited. One reason for this is the difficulty adapting the BERT architecture to the partially observable Markov decision process present in VLN, requiring history-dependent attention and decision making. In this paper we propose a recurrent BERT model that is time-aware for use in VLN. Specifically, we equip the BERT model with a recurrent function that maintains cross-modal state information for the agent. Through extensive experiments on R2R and REVERIE we demonstrate that our model can replace more complex encoder-decoder models to achieve state-of-the-art results. Moreover, our approach can be generalised to other transformer-based architectures, supports pre-training, and is capable of solving navigation and referring expression tasks simultaneously.
Description
Keywords
Citation
Collections
Source
Proceedings of the IEEE Computer Society Conference on Computer Vision and Pattern Recognition
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
2099-12-31