Cultural advice

The Australian National University acknowledges, celebrates and pays our respects to the Ngunnawal and Ngambri people of the Canberra region and to all First Nations Australians on whose traditional lands we meet and work, and whose cultures are among the oldest continuing cultures in human history.

Aboriginal and Torres Strait Islander peoples are advised that ANU Library collections may include images, names, voices, and other representations of deceased persons.

Material in the collection may contain terms, language or views that reflect the period in which the item was created and may be considered inappropriate today.

Transferring Cross-Domain Knowledge for Video Sign Language Recognition

dc.contributor.authorLi, Dongxu
dc.contributor.authorYu, Xin
dc.contributor.authorXu, Chenchen
dc.contributor.authorPetersson, Lars
dc.contributor.authorLi, Hongdong
dc.coverage.spatialUnited States
dc.date.accessioned2024-01-21T23:05:54Z
dc.date.createdJune 14-19 2020
dc.date.issued2020
dc.date.updated2022-10-02T07:17:06Z
dc.description.abstractWord-level sign language recognition (WSLR) is a fundamental task in sign language interpretation. It requires models to recognize isolated sign words from videos. However, annotating WSLR data needs expert knowledge, thus limiting WSLR dataset acquisition. On the contrary, there are abundant subtitled sign news videos on the internet. Since these videos have no word-level annotation and exhibit a large domain gap from isolated signs, they cannot be directly used for training WSLR models. We observe that despite the existence of a large domain gap, isolated and news signs share the same visual concepts, such as hand gestures and body movements. Motivated by this observation, we propose a novel method that learns domain-invariant visual concepts and fertilizes WSLR models by transferring knowledge of subtitled news sign to them. To this end, we extract news signs using a base WSLR model, and then design a classifier jointly trained on news and isolated signs to coarsely align these two domain features. In order to learn domain-invariant features within each class and suppress domain-specific features, our method further resorts to an external memory to store the class centroids of the aligned news signs. We then design a temporal attention based on the learnt descriptor to improve recognition performance. Experimental results on standard WSLR datasets show that our method outperforms previous state-of-the-art methods significantly. We also demonstrate the effectiveness of our method on automatically localizing signs from sign news, achieving 28.1 for AP@0.5.en_AU
dc.format.mimetypeapplication/pdfen_AU
dc.identifier.isbn978-172819360-1en_AU
dc.identifier.urihttp://hdl.handle.net/1885/311657
dc.language.isoen_AUen_AU
dc.publisherIEEE Computer Societyen_AU
dc.relation.ispartofseries2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2020en_AU
dc.rights© 2020 IEEEen_AU
dc.source2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, CVPRW 2020en_AU
dc.titleTransferring Cross-Domain Knowledge for Video Sign Language Recognitionen_AU
dc.typeConference paperen_AU
local.bibliographicCitation.lastpage6213en_AU
local.bibliographicCitation.startpage6204en_AU
local.contributor.affiliationLi, Dongxu, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationYu, Xin, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationXu, Chenchen, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationPetersson, Lars, College of Engineering and Computer Science, ANUen_AU
local.contributor.affiliationLi, Hongdong, College of Engineering and Computer Science, ANUen_AU
local.contributor.authoruidLi, Dongxu, u5886609en_AU
local.contributor.authoruidYu, Xin, u5819038en_AU
local.contributor.authoruidXu, Chenchen, u5545924en_AU
local.contributor.authoruidPetersson, Lars, u4048690en_AU
local.contributor.authoruidLi, Hongdong, u4056952en_AU
local.description.embargo2099-12-31
local.description.notesImported from ARIESen_AU
local.description.refereedYes
local.identifier.absfor470407 - Language documentation and descriptionen_AU
local.identifier.absfor460309 - Video processingen_AU
local.identifier.absfor460208 - Natural language processingen_AU
local.identifier.ariespublicationa383154xPUB14012en_AU
local.identifier.doi10.1109/CVPR42600.2020.00624en_AU
local.identifier.scopusID2-s2.0-85094646882
local.identifier.thomsonIDWOS:000620679506048
local.publisher.urlhttps://www.ieee.org/en_AU
local.type.statusPublished Versionen_AU

Downloads

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Transferring_Cross-Domain_Knowledge_for_Video_Sign_Language_Recognition.pdf
Size:
1.11 MB
Format:
Adobe Portable Document Format
Description: