Guided open vocabulary image captioning with constrained beam search
| dc.contributor.author | Anderson, Peter | |
| dc.contributor.author | Fernando, Basura | |
| dc.contributor.author | Johnson, Mark | |
| dc.contributor.author | Gould, Stephen | |
| dc.contributor.editor | Martha Palmer | |
| dc.contributor.editor | Rebecca Hwa | |
| dc.contributor.editor | Sebastian Riedel | |
| dc.coverage.spatial | Copenhagen, Denmark | |
| dc.date.accessioned | 2023-07-20T00:51:50Z | |
| dc.date.available | 2023-07-20T00:51:50Z | |
| dc.date.created | September 7-11 2017 | |
| dc.date.issued | 2017 | |
| dc.date.updated | 2022-05-22T08:15:50Z | |
| dc.description.abstract | Existing image captioning models do not generalize well to out-of-domain images containing novel scenes or objects. This limitation severely hinders the use of these models in real world applications dealing with images in the wild. We address this problem using a flexible approach that enables existing deep captioning architectures to take advantage of image taggers at test time, without re-training. Our method uses constrained beam search to force the inclusion of selected tag words in the output, and fixed, pretrained word embeddings to facilitate vocabulary expansion to previously unseen tag words. Using this approach we achieve state of the art results for out-of-domain captioning on MSCOCO (and improved results for in-domain captioning). Perhaps surprisingly, our results significantly outperform approaches that incorporate the same tag predictions into the learning algorithm. We also show that we can significantly improve the quality of generated ImageNet captions by leveraging ground-truth labels. | en_AU |
| dc.description.sponsorship | This research is supported by an Australian Government Research Training Program (RTP) Scholarship and by the Australian Research Council Centre of Excellence for Robotic Vision (project number CE140100016). | en_AU |
| dc.format.mimetype | application/pdf | en_AU |
| dc.identifier.isbn | 978-1-945626-83-8 | en_AU |
| dc.identifier.uri | http://hdl.handle.net/1885/294445 | |
| dc.language.iso | en_AU | en_AU |
| dc.provenance | https://aclanthology.org/faq/..."The ACL materials that are hosted in the Anthology are licensed to the general public under a liberal usage policy that allows unlimited reproduction, distribution and hosting of materials on any other website or medium, for non-commercial purposes. Prior to 2016, all ACL materials are licensed using the Creative Commons 3.0 BY-NC-SA (Attribution, Non-Commercial, Share-Alike) license. As of 2016, this policy has been relaxed further, and all subsequent materials are available to the general public on the terms of the Creative Commons 4.0 BY (Attribution) license; this means both commercial and non-commercial use is explicitly licensed to all." from the publisher site (as at 20.07.2023) | en_AU |
| dc.publisher | Association for Computational Linguistics | en_AU |
| dc.relation | http://purl.org/au-research/grants/arc/CE140100016 | en_AU |
| dc.relation.ispartofseries | Conference on Empirical Methods in Natural Language Processing, EMNLP2017 | en_AU |
| dc.rights | © 2017 Association for Computational Linguistics | en_AU |
| dc.rights.license | Creative Commons Attribution 4.0 International License | en_AU |
| dc.rights.uri | https://creativecommons.org/licenses/by/4.0/ | en_AU |
| dc.source | Proceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP2017 | en_AU |
| dc.source.uri | https://aclanthology.org/D17-1098/ | en_AU |
| dc.title | Guided open vocabulary image captioning with constrained beam search | en_AU |
| dc.type | Conference paper | en_AU |
| dcterms.accessRights | Open Access | en_AU |
| local.bibliographicCitation.lastpage | 945 | en_AU |
| local.bibliographicCitation.startpage | 936 | en_AU |
| local.contributor.affiliation | Anderson, Peter, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.affiliation | Fernando, Basura, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.affiliation | Johnson, Mark, Macquarie University | en_AU |
| local.contributor.affiliation | Gould, Stephen, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.authoruid | Anderson, Peter, u4304616 | en_AU |
| local.contributor.authoruid | Fernando, Basura, u1000328 | en_AU |
| local.contributor.authoruid | Gould, Stephen, u4971180 | en_AU |
| local.description.notes | Imported from ARIES | en_AU |
| local.description.refereed | Yes | |
| local.identifier.absfor | 460200 - Artificial intelligence | en_AU |
| local.identifier.ariespublication | a383154xPUB12245 | en_AU |
| local.identifier.doi | 10.18653/v1/D17-1098 | en_AU |
| local.publisher.url | https://aclanthology.org/D17-1098/ | en_AU |
| local.type.status | Published Version | en_AU |
Downloads
Original bundle
1 - 1 of 1
Loading...
- Name:
- Guided Open Vocabulary Image Captioning.pdf
- Size:
- 1.81 MB
- Format:
- Adobe Portable Document Format
- Description: