This deep learning-based scene reconstruction technology takes monocular RGB image sequences and camera pose data as input, generates a 3D feature volume through a fusion of CNN and GRU, and predicts it as a TSDF volume to reconstruct a dense 3D mesh.
Existing monocular RGB-based 3D reconstruction technologies often suffer from high dependency on depth map quality, high computational costs during real-time reconstruction, and low reconstruction completeness, making precise scene representation difficult.
This technology optimizes reconstruction performance and efficiency by combining keyframe selection, local fragment segmentation, a feature extraction network, a 3D CNN and GRU fusion unit, and a refinement network. It can be applied to autonomous driving, AR/VR content creation, and robotic spatial awareness, enabling the acquisition of precise 3D spaces using only a camera, without the need for depth sensors.
This invention was developed with support from the Ministry of Science and ICT for the development of robust pose estimation and 3D environment reconstruction algorithms through the fusion of event cameras, physical sensors, and deep learning in extreme environments.
N/A