Towards Efficient, Accurate and Self-trained Multi-view Stereopsis
Abstract
Reconstructing 3D models from multi-view images is a fundamental problem in computer vision, with applications in fields such as virtual reality, robotics, archaeology, and cultural heritage preservation. Multi-view Stereo (MVS) is a key step in 3D reconstruction pipelines, which uses multiple posed images of an object to estimate its detailed geometry. Traditional optimisation-based methods minimise hand-crafted photometric matching costs between multiple views to infer 3D geometry. Recent deep learning-based MVS networks learn discriminative feature and context-aware cost aggregation to achieve greater accuracy and efficiency. While these methods have shown promising results, critical challenges remain. In this thesis, we present a series of methods aimed at tackling the three primary challenges in this area: (1) accuracy, (2) efficiency, and (3) training strategy. These methods represent a line of techniques that have achieved state-of-the-art performance, offering a promising outlook for the future of Multi-view Stereo.
First, we tackle the efficiency of learning-based multi-view stereo networks by a cost volume pyramid pipeline. We build a cost volume pyramid in a coarse-to-fine manner which leads to a compact, lightweight network allowing inferring high-resolution depth maps at a lower computational cost to achieve better reconstruction results.
Secondly, we mitigate the need for ground-truth training data by proposing a self-supervised learning framework for multi-view stereo that exploits pseudo labels from the input data. We start by learning to estimate depth maps as initial pseudo labels and then refine the initial pseudo labels using a carefully designed pipeline. We use these high-quality pseudo labels as the supervision signal to train the network and improve, iteratively, its performance by self-training.
Thirdly, we address the issue of reconstruction accuracy for small objects and boundary regions by utilising non-parametric depth distribution modelling to handle pixels exhibiting both unimodal and multi-modal distributions. Our approach is designed to output multiple depth hypotheses at the coarser levels and to maintain rigid spatial relationships using a sparse cost aggregation network. Additionally, we introduce a novel application of our aforementioned depth distribution modelling techniques for enhancing the perception capabilities of autonomous vehicles in Bird Eye's View (BEV) space. We employ parametric depth distribution modelling to transform image features into BEV space and estimate BEV visibility, which holds great potential for improving the safety and efficiency of autonomous driving systems.
In summary, the goal of this thesis is to enable vision systems to effectively reconstruct 3D structure of real-life scenarios. We propose several novel ideas that push the boundaries of current state-of-the-arts solutions for Multi-view Stereo.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description
Thesis Material