Zhong, YiranDai, YuchaoLi, Hongdong2020-02-12August 20-1051-4651http://hdl.handle.net/1885/201664This paper is concerned with the problem of how to better exploit 3D geometric information for dense semantic image labeling. Existing methods often treat the available 3D geometry information (e.g., 3D depth-map) simply as an additional image channel besides the R-G-B color channels, and apply the same technique for RGB image labeling. In this paper, we demonstrate that directly performing 3D convolution in the framework of a residual connected 3D voxel top-down modulation network can lead to superior results. Specifically, we propose a 3D semantic labeling method to label outdoor street scenes whenever a dense depth map is available. Experiments on the 'Synthia' and 'Cityscape' datasets show our method outperforms the state-of-the-art methods, suggesting such a simple 3D representation is effective in incorporating 3D geometric information.We gratefully acknowledge the support of NVIDIA Corporation with donation of TITAN Xp GPU used for this research, as well a NVIDIA Drive-PX2 platform for an autonomous driving project. YZ’s PhD scholarship is funded by CSIRO Data61. Y. Dai was supported in part by National 1000 Young Talents Plan of China, Natural Science Foundation of China (61420106007, 61671387), and ARC grant (DE140100180). H. Li’s work is funded in part by Australia ARC Centre of Excellence for Robotic Vision (CE140100016).7 pagesapplication/pdfen-AU© 2018 IEEEThree-dimensional displays, Semantics, Convolution, Labeling, Feature extraction, Two dimensional displays, Geometry3D Geometry-Aware Semantic Labeling of Outdoor Street Scenes2018-11-2910.1109/ICPR.2018.85453782019-11-25