Multi-view Inverse Rendering: Reconstructing 3D Shape and Reflectance with a Camera and Light Source
Abstract
Rendering in computer graphics refers to photo-realistic image synthesis from intrinsic properties of objects, such as their geometry and material's reflectance, under given lighting and camera configurations. In contrast, inverse rendering is regarded as a computer vision task, whose aim is to recover the intrinsic properties of the scene from a set of captured images. The end-product is typically a geometric and photometric virtualization of object that can be rendered under novel view point and/or illumination.
Inverse rendering is an important task in computer vision research, and can be regarded as a broad generalization of several well known and well researched topics, such as Multi-View-Stereo, Shape-from-Shading, Intrinsic-Image-Decomposition, Photometric-Stereo and Novel-View-Synthesis. Consequently, many downstream applications exist. These include object 3D reconstruction, human body and face scan, material virtualization and classification, and mixed reality.
The problem itself, however, is as challenging as it is rewarding. In a generic setup, both multi-view geometric and photometric constraints need to be jointly addressed, so that not only the 3D geometry but also the appearance can be faithfully recovered. Traditional stereo reconstruction methods only tackle one side of this problem: Multi-View-Stereo assumes static photometric appearance and only solves for geometry, while Photometric-Stereo assumes static camera pose to circumvent geometric correspondence matching. The challenge of full inverse rendering arises from the fact that establishing multi-view correspondences is difficult when the appearance is view-dependent, and even more so under varying illumination conditions. On the other hand, reconstructing the material reflectance depends on good multi-view correspondences, resulting in a chicken-egg problem. Moreover, the underlying mathematical problem is a highly non-convex one. Therefore care must be taken in its parameterization and optimization.
To overcome above issues, we propose three solutions for full inverse rendering using a moving camera and moving light source. This configuration allows images to be captured under different view points and lighting conditions, enabling sufficient geometric and photometric constraints on object intrinsics. All three solutions recover full 3D shape and its generic reflectance model. In the first pipeline, we show that the inverse rendering problem becomes well-posed with multi-view images and a moving light source. The underlying optimization problem, while being still highly non-convex, can be solved in a globally optimal manner by using appropriate relaxation and randomized search algorithms. In the second work, we use a neural ordinary differential equation for topology-preserving 3D shape and reflectance map parameterization. As a result, we were able to solve the same non-convex problem but by standard gradient descent when guaranteeing the integrity of neural representations. These two works, however, are limited to a darkroom environment without any ambient light contamination. In the third and final work, we remove this limitation by incorporating an ambient reflection neural light field that is independent from the moving light source, thereby achieving in-the-wild inverse rendering. All three methods are workable with a hand-held device, and can recover detailed geometry and highly specular reflectance.
Description
Keywords
Citation
Collections
Source
Type
Book Title
Entity type
Access Statement
License Rights
Restricted until
Downloads
File
Description
Thesis Material