Welcome to Journal of Graphics

Journal of Graphics ›› 2026, Vol. 47 ›› Issue (4): 704-713.DOI: 10.11996/JG.j.2095-302X.2026040704

• Image Processing and Computer Vision • Previous Articles     Next Articles

Few-shot 3D Gaussian splatting based on monocular depth ambiguity-aware estimation

WANG Ziwei, WANG Lutao(), LI Antong, SHEN Yan   

  1. College of Computer Science and technology, Chengdu University of Information Technology, Chengdu Sichuan 610225, China
  • Received:2025-11-27 Accepted:2026-03-22 Online:2026-08-31 Published:2026-08-31
  • Contact: WANG Lutao
  • Supported by:
    National Natural Science Foundation of China(62172061)

Abstract:

After Neural Radiance Fields (NeRF), which predominantly use coordinate-based models to map spatial coordinates to pixel values, 3D Gaussian Splatting (3DGS) has emerged as an important breakthrough technique in the realm of three-dimensional reconstruction and novel view synthesis. By transforming multi-view images into millions of learnable 3D Gaussians to model an explicit scene representation, which introduces unprecedented levels of editability, and using differentiable rendering, 3DGS achieves near real-time view rendering. However, 3DGS is prone to overfitting the training views when a small number of images are available due to Gaussian splats’ local nature. Moreover, 3DGS optimizes independent splats only under multi-view color supervision, without global geometric structure. This problem becomes more pronounced as the number of images used for 3D scene optimization decreases, because sufficient images that can offer global geometric cues are unavailable. As a result, reconstructed scenes under sparse-view input often suffer from problems such as holes and floating artifacts. To address this issue, an algorithm of few-shot 3D Gaussian splatting based on monocular depth ambiguity-aware estimation was proposed to improve the quality of scene reconstruction and the effect of novel view synthesis, thereby further expanding the 3DGS application scenario. For sparse-input settings, Conditional Implicit Maximum Likelihood Estimation (CIMLE) was used to learn the multimodal distribution of the depth estimation to avoid the unimodal distribution that may be caused by the conditional GAN mode collapse. This monocular depth ambiguity-aware technique was used to extract a multi-modal depth dense point cloud for initialization of Gaussian primitives, and to introduce global geometric cues for scene reconstruction. Then, a space-carving loss was designed, along with color-based photometric loss optimization to jointly supervise and optimize the scene structure, thereby resolving the inherent uncertainty and ambiguity retained by the spatial positions of Gaussian primitives in the initialization phase. The multi-modal depth distribution was used to constrain the optimization of Gaussian primitives, with the goal of finding a globally consistent subset of patterns captured from the monocular depth ambiguity-aware distribution of each view. In this way, information from different views was fused together, since the inherent ambiguity can only be resolved through multi-view information. This allowed a common shape consistent across all views and the object surface to be “snapped” together. This approach alleviated scene overfitting under sparse views and effectively improved the quality of scene reconstruction. The experimental results showed that the average PSNR performance of the proposed three-dimensional Gaussian reconstruction algorithm was 11%-54% higher than that of the same period algorithm. At the same time, training time was greatly reduced compared with the NeRF-based algorithm, and improved scene reconstruction and novel-view-synthesis quality was achieved.

Key words: three-dimensional reconstruction, sparse input, novel view synthesis, 3D Gaussian splatting, monocular depth estimation

CLC Number: