Multiview Detection with Shadow Transformer (and View-Coherent Data Augmentation)
| dc.contributor.author | Hou, Yunzhong | |
| dc.contributor.author | Zheng, Liang | |
| dc.coverage.spatial | Virtual Event China | |
| dc.date.accessioned | 2024-01-29T23:17:32Z | |
| dc.date.created | October 20 - 24, 2021 | |
| dc.date.issued | 2021 | |
| dc.date.updated | 2022-10-02T07:18:44Z | |
| dc.description.abstract | Multiview detection incorporates multiple camera views to deal with occlusions, and its central problem is multiview aggregation. Given feature map projections from multiple views onto a common ground plane, the state-of-the-art method addresses this problem via convolution, which applies the same calculation regardless of object locations. However, such translation-invariant behaviors might not be the best choice, as object features undergo various projection distortions according to their positions and cameras. In this paper, we propose a novel multiview detector, MVDeTr, that adopts a newly introduced shadow transformer to aggregate multiview information. Unlike convolutions, shadow transformer attends differently at different positions and cameras to deal with various shadow-like distortions. We propose an effective training scheme that includes a new view-coherent data augmentation method, which applies random augmentations while maintaining multiview consistency. On two multiview detection benchmarks, we report new state-of-the-art accuracy with the proposed system. Code is available at https://github.com/hou-yz/MVDeTr. | en_AU |
| dc.description.sponsorship | This work was supported by the ARC Discovery Early Career Researcher Award (DE200101283) and the ARC Discovery Project (DP210102801). | en_AU |
| dc.format.mimetype | application/pdf | en_AU |
| dc.identifier.isbn | 978-1-4503-8651-7 | en_AU |
| dc.identifier.uri | http://hdl.handle.net/1885/312408 | |
| dc.language.iso | en_AU | en_AU |
| dc.publisher | Association for Computing Machinery (ACM) | en_AU |
| dc.relation | http://purl.org/au-research/grants/arc/DE200101283 | en_AU |
| dc.relation | http://purl.org/au-research/grants/arc/DP210102801 | en_AU |
| dc.relation.ispartofseries | MM '21: ACM Multimedia Conference | en_AU |
| dc.rights | © 2021 Copyright held by the owner/author(s). Publication rights licensed to ACM | en_AU |
| dc.source | Proceedings of the 29th ACM International Conference on Multimedia | en_AU |
| dc.subject | multiview detection | en_AU |
| dc.subject | transformer | en_AU |
| dc.subject | data augmentation | en_AU |
| dc.title | Multiview Detection with Shadow Transformer (and View-Coherent Data Augmentation) | en_AU |
| dc.type | Conference paper | en_AU |
| local.bibliographicCitation.lastpage | 1682 | en_AU |
| local.bibliographicCitation.startpage | 1673 | en_AU |
| local.contributor.affiliation | Hou, Yunzhong, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.affiliation | Zheng, Liang, College of Engineering and Computer Science, ANU | en_AU |
| local.contributor.authoruid | Hou, Yunzhong, u6852178 | en_AU |
| local.contributor.authoruid | Zheng, Liang, u1064892 | en_AU |
| local.description.embargo | 2099-12-31 | |
| local.description.notes | Imported from ARIES | en_AU |
| local.description.refereed | Yes | |
| local.identifier.absfor | 461103 - Deep learning | en_AU |
| local.identifier.absfor | 460304 - Computer vision | en_AU |
| local.identifier.ariespublication | a383154xPUB24134 | en_AU |
| local.identifier.doi | 10.1145/3474085.3475310 | en_AU |
| local.identifier.scopusID | 2-s2.0-85119338914 | |
| local.publisher.url | https://dl.acm.org/doi/10.1145/3474085.3475310 | en_AU |
| local.type.status | Published Version | en_AU |
Downloads
Original bundle
1 - 1 of 1
Loading...
- Name:
- Multiview Detection with Shadow Transformer.pdf
- Size:
- 6.73 MB
- Format:
- Adobe Portable Document Format
- Description: