NeRFifyMesh: Optimizing Neural Radiance Fields from Textured Meshes for Robotics Scene Building
NeRFifyMesh: Optimizing Neural Radiance Fields from Textured Meshes for Robotics Scene Building
Nillan Nimal1, Mahboubeh Asadi2, Sajad Saeedi3
1University of Toronto, 2Toronto Metropolitan University,3University College London
In robotics, scene representation plays a pivotal role in understanding and interacting with the environment. The advent of Neural Radiance Fields (NeRF) and its variants, as a novel representation, has opened a new frontier of research. In applications such as semantic mapping and simulation, roboticists aim to build scenes using multiple NeRF models, each representing an object. While extensive datasets of 3D mesh models already exist, there is an urgent need to develop tools to convert these assets to NeRF models for rapid algorithm development and testing. This paper presents a new pipeline for converting existing mesh models to NeRF representations by artificially generating a ground truth point-based radiance field through sampling mesh geometry and texture. This approach alleviates the need for camera-based sampling or rendering multi-view images of the original mesh to train the NeRF model. Extensive benchmarking demonstrates that our method yields comparable rendering quality to the baselines. Additionally, the application of this representation is shown by constructing unified NeRF scenes and performing collision simulations with extracted geometry.
We propose NeRFifyMesh, a novel pipeline that directly optimizes a NeRF from textured mesh models. Our pipeline removes the rendering optimization step from the NeRF training process and alleviates the need for high-quality ground truth images and camera poses of the mesh model. We achieve this by sampling the mesh to generate a discrete point-based radiance field with the properties of emitted radiance (r,g,b), binary opacity (α), normals (n) , and their significance, wc, w𝞪, and wn, respectively. We then proceed to train a neural network directly on this discrete field. NeRFifyMesh supports scene rendering, normal prediction, surface recovery, and scene composition for robotics applications.
NeRFifyMesh decouples the prediction of opacity, color and normals into three separate networks. Input points x are queried across multi-resolution surface voxel grids with hash-encoded features. Interpolated features are decoded with network fψ and combined with the original input to predict opacity through network fζ . Positional encoding of the input is provided to networks fθ and fϕ to predict colors and normals, respectively.
Shown below is a visual comparison of mesh scene rendering results for NeRFifyMesh, Instant-NGP, and NeRF across objects from the Google Scanned Objects (GSO), and Objaverse datasets. NeRF and Instant-NGP were trained using multi-view images of the mesh, while NeRFifyMesh was trained directly from mesh data.
Since NeRFifyMesh predicts point normals, it supports relighting through the Blinn-Phong illumination model. The renders below demonstrate a Horse-Conch scene that is rendered under different lighting conditions by varying ambient, diffuse and, specular lighting parameters: (a) original scene without altered lighting, (b) ambient only, (c) ambient + diffuse, (d) ambient + diffuse + specular at low intensity, and (e) the same with high specular intensity.
Following the approach of NeRF2Real, we create dynamic simulations from our models by leveraging the PyBullet physics engine. The engine is initialized with meshes extracted from trained NeRFifyMesh models to form the collision representation. For each rendered simulation frame, object poses from the physics engine are utilized to perform volume rendering and the rendered objects are composited to form the global scene.
Simulations can be formed from static NeRF scenes and our models. This video shows a simulation of dropping a football (Poly Haven dataset) from a height of 0.78m onto a table within the NeRF Garden scene. This experiment shows how our representation can facilitate combined simulations of static NeRF scenes with dynamic objects formed from meshes for robotics applications.
The video below shows a collision simulation composed of two objects situated on a ground plane. We compare our results against a simulation created using Instant-NGP, where collision geometry is extracted with Marching Cubes. The ground-truth simulation uses the original mesh for collision geometry and mesh-based rendering for visualization.
Ground Truth
Instant-NGP
NeRFifyMesh
If you have any questions, feel free to reach out to us at the following email us at: nillan.nimal@mail.utoronto.ca