Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Guided open vocabulary image captioning with constrained beam search

dc.contributor.authorAnderson, Peter
dc.contributor.authorFernando, Basura
dc.contributor.authorJohnson, Mark
dc.contributor.authorGould, Stephen
dc.contributor.editorMartha Palmer
dc.contributor.editorRebecca Hwa
dc.contributor.editorSebastian Riedel
dc.coverage.spatialCopenhagen, Denmark
dc.date.accessioned2023-07-20T00:51:50Z
dc.date.available2023-07-20T00:51:50Z
dc.date.createdSeptember 7-11 2017
dc.date.issued2017
dc.date.updated2022-05-22T08:15:50Z
dc.description.abstractExisting image captioning models do not generalize well to out-of-domain images containing novel scenes or objects. This limitation severely hinders the use of these models in real world applications dealing with images in the wild. We address this problem using a flexible approach that enables existing deep captioning architectures to take advantage of image taggers at test time, without re-training. Our method uses constrained beam search to force the inclusion of selected tag words in the output, and fixed, pretrained word embeddings to facilitate vocabulary expansion to previously unseen tag words. Using this approach we achieve state of the art results for out-of-domain captioning on MSCOCO (and improved results for in-domain captioning). Perhaps surprisingly, our results significantly outperform approaches that incorporate the same tag predictions into the learning algorithm. We also show that we can significantly improve the quality of generated ImageNet captions by leveraging ground-truth labels.en_AU
dc.description.sponsorshipThis research is supported by an Australian Government Research Training Program (RTP) Scholarship and by the Australian Research Council Centre of Excellence for Robotic Vision (project number CE140100016).en_AU
dc.format.mimetypeapplication/pdfen_AU
dc.identifier.isbn978-1-945626-83-8en_AU
dc.identifier.urihttp://hdl.handle.net/1885/294445
dc.language.isoen_AUen_AU
dc.provenancehttps://aclanthology.org/faq/..."The ACL materials that are hosted in the Anthology are licensed to the general public under a liberal usage policy that allows unlimited reproduction, distribution and hosting of materials on any other website or medium, for non-commercial purposes. Prior to 2016, all ACL materials are licensed using the Creative Commons 3.0 BY-NC-SA (Attribution, Non-Commercial, Share-Alike) license. As of 2016, this policy has been relaxed further, and all subsequent materials are available to the general public on the terms of the Creative Commons 4.0 BY (Attribution) license; this means both commercial and non-commercial use is explicitly licensed to all." from the publisher site (as at 20.07.2023)en_AU
dc.publisherAssociation for Computational Linguisticsen_AU
dc.relationhttp://purl.org/au-research/grants/arc/CE140100016en_AU
dc.relation.ispartofseriesConference on Empirical Methods in Natural Language Processing, EMNLP2017en_AU
dc.rights© 2017 Association for Computational Linguisticsen_AU
dc.rights.licenseCreative Commons Attribution 4.0 International Licenseen_AU
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/en_AU
dc.sourceProceedings of the Conference on Empirical Methods in Natural Language Processing, EMNLP2017en_AU
dc.source.urihttps://aclanthology.org/D17-1098/en_AU
dc.titleGuided open vocabulary image captioning with constrained beam searchen_AU
dc.typeConference paperen_AU
dcterms.accessRightsOpen Accessen_AU
local.bibliographicCitation.lastpage945en_AU
local.bibliographicCitation.startpage936en_AU
local.contributor.affiliationAnderson, Peter, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationFernando, Basura, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationJohnson, Mark, Macquarie Universityen_AU
local.contributor.affiliationGould, Stephen, College of Engineering and Computer Science, ANUen_AU
local.contributor.authoruidAnderson, Peter, u4304616en_AU
local.contributor.authoruidFernando, Basura, u1000328en_AU
local.contributor.authoruidGould, Stephen, u4971180en_AU
local.description.notesImported from ARIESen_AU
local.description.refereedYes
local.identifier.absfor460200 - Artificial intelligenceen_AU
local.identifier.ariespublicationa383154xPUB12245en_AU
local.identifier.doi10.18653/v1/D17-1098en_AU
local.publisher.urlhttps://aclanthology.org/D17-1098/en_AU
local.type.statusPublished Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Guided Open Vocabulary Image Captioning.pdf
Size:
1.81 MB
Format:
Adobe Portable Document Format
Description: