Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction

Wei-Teng (Wayne) Chu* · Yashasvini Gopalan* · Changju Yuan*
Stanford University  ·  *Equal contribution
CS231N: Deep Learning for Computer Vision — Final Project
Teaser / pipeline overview
Given sparse forward-facing RGB inputs, we estimate a monocular depth prior with Depth Anything V2 and align it to metric depth via scale-shift fitting. An RGB-only baseline yields per-pixel photometric error, which builds a reliability mask that keeps low-error regions. The aligned prior is applied only to reliable pixels through a masked depth loss while the model is jointly optimized with the photometric loss.

Abstract

Sparse-view neural reconstruction is challenging in outdoor driving scenes, where cameras move along a narrow forward-facing trajectory and provide limited multi-view overlap. Monocular depth estimators can provide dense geometric priors, but their predictions are noisy and not uniformly reliable across image regions. We use Depth Anything V2 (DA-V2) as a dense monocular depth prior, align its predictions to metric depth via per-image scale-shift fitting on sparse anchors (LiDAR on KITTI, COLMAP on Bicycle), and apply depth supervision selectively through photometric masks generated from an RGB-only baseline. We evaluate this strategy on two representative scene representations: Mip-NeRF-360 and Splatfacto. Since the mask modifies an existing depth loss rather than introducing one, we compare against global (unmasked) supervision as the primary baseline. Across KITTI sequences 00/02/05, the mask yields +0.44 to +0.70 dB PSNR for Splatfacto over global supervision at tied or better RMSE, and no change on Mip-NeRF-360 — the mask primarily enhances rendering fidelity rather than geometry. Matched-ratio ablations confirm the gains come from selecting reliable low-error regions rather than simply reducing the number of depth-supervised pixels. Without any ground-truth depth, our mask also outperforms three LiDAR-supervised baselines. On the object-centric Mip-NeRF-360 Bicycle scene, depth supervision improves geometry but hurts rendering quality when multi-view coverage is already strong.

Method

Scale-Aligned Monocular Depth Prior

DA-V2 predicts relative depth, so we align each prediction to a metric reference before training. For KITTI the reference is valid LiDAR depth; for the Bicycle scene it is sparse COLMAP points. Given a monocular prediction \(d_m(u)\) and reference depth \(d_r(u)\) over valid anchor pixels \(u \in \Omega\), we solve a per-image least-squares scale-shift fit:

\[ s^{*},\, t^{*} = \operatorname*{arg\,min}_{s,t} \sum_{u \in \Omega} \big( s\, d_m(u) + t - d_r(u) \big)^2 \]

The aligned prior is \(\hat{d}(u) = s^{*} d_m(u) + t^{*}\), clipped to the 80 m evaluation range on KITTI.

Photometric Reliability Mask

Rather than supervising every pixel, we only supervise where the RGB-only baseline is already reliable. From a baseline render \(\hat{I}\) and the ground-truth image \(I\), we compute a per-pixel photometric error and threshold it at \(\tau\):

\[ e(u) = \frac{1}{3} \sum_{c \in \{R,G,B\}} \big| \hat{I}_c(u) - I_c(u) \big|, \qquad M_\tau(u) = \mathbb{1}\!\left[\, e(u) < \tau \,\right] \]

High photometric error flags unreliable regions — dynamic objects, reflective surfaces, occlusion boundaries, and sky — where a depth loss would push the model toward wrong geometry. Low-error pixels are already consistent with the RGB observations, so the aligned prior acts as a stabilizing regularizer. The mask is computed once from the RGB-only baseline and held fixed throughout depth-supervised training.

Masked Depth Supervision Objective

The effective supervision mask combines the photometric mask with the depth-validity mask \(D(u)\): \(M_\text{eff} = M_\tau \wedge D\). The depth loss is gated by \(M_\text{eff}\), while the RGB loss is evaluated over the full image:

\[ \mathcal{L} = \mathcal{L}_\text{rgb} + \lambda_\text{depth}\, \mathcal{L}_\text{depth}, \qquad \mathcal{L}_\text{depth} = \frac{1}{N} \sum_{u} M_\text{eff}(u) \big( \hat{d}(u) - \hat{d}_\text{prior}(u) \big)^2 \]

Setting \(\tau=1.0\) gives \(M_\tau \equiv 1\) — global (unmasked) depth supervision. This condition, not the depth-free \(\lambda_\text{depth}=0\) baseline, is the primary baseline for judging the mask's contribution: it isolates the effect of where depth is supervised from the effect of whether depth is supervised at all.

Dataset and mask-generation pipeline
Training views are subsampled from KITTI Odometry to simulate sparse-view capture. Fixed photometric masks are generated from baseline renderings by thresholding the photometric error; DA-V2 predictions are scale-shift aligned to produce metric depth priors.

Depth Prior Quality

Before using DA-V2 as supervision, we evaluate the aligned depth against valid KITTI LiDAR pixels. The mean absolute error of alignment is 4.22 m. Sweeping the number of LiDAR anchors used for scale-shift fitting shows the 2-DOF optimization saturates long before typical LiDAR density is reached — alignment MAE is stable from ~95k anchors down to 500 per frame, degrading only under extreme sparsity. The 4.22 m error floor reflects DA-V2's intrinsic local structural limitations rather than calibration artifacts.

Anchors / frameAlignment MAE ↓ (m)
~95,0004.22
5,0004.22
5004.22
1004.29
504.29
204.51

Photometric Error as a Reliability Proxy

Because photometric and depth errors are measured against different references, low photometric error does not by construction imply low depth error. We validate the mask directly against GT LiDAR. Depth RMSE is 34–39% lower inside the mask at every threshold, and inside-mask RMSE increases monotonically with τ (r=1.0). Pixel-level differences are highly significant (p<10−50); the frame-level paired t-test gives p=0.047 (n=8).

τInside RMSE ↓ (m)Outside RMSE ↓ (m)
0.146.309.58
0.166.329.82
0.18 (ours)6.3310.01
0.206.3610.24
0.226.3810.45
1.00 (global)6.43—

Results

Splatfacto — KITTISeq02 every-2

Splatfacto gains clearly from monocular depth supervision. Using any DA-V2 depth prior — even without a mask — improves PSNR (14.903 → 15.494 for global) and cuts RMSE by ~80% (0.542 → 0.101). Masking the depth loss at τ=0.18 adds a further +0.44 dB PSNR at tied RMSE. In short: the RMSE drop comes from using a depth prior at all, and masking is what improves rendering fidelity. Over three seeds, masked 15.30±0.57 vs global 15.07±0.38 dB (+0.23 dB mean); the −80% RMSE improvement over RGB-only is consistent across seeds.

Setting (τ, λ)PSNR ↑SSIM ↑LPIPS ↓RMSE ↓
RGB-only14.9030.4330.4460.542
0.16, 0.1015.4520.4540.4340.106
0.18, 0.10 (ours)15.9320.4770.4080.100
0.18, 0.1515.5880.4590.4360.096
1.00, 0.10 (global)15.4940.4480.4340.101

Matched-Ratio Mask Ablation

To test whether the benefit comes from selecting reliable pixels or simply using fewer pixels, we compare against high-error and random masks with the identical per-frame pixel budget. The low-error mask wins on every metric — the improvement is not from fewer supervised pixels.

Mask type (λ=0.10)PSNR ↑SSIM ↑LPIPS ↓RMSE ↓
High-error, matched14.9320.4370.4550.111
Random, matched (3 seeds)15.0360.4420.4560.109
Low-error, τ=0.18 (ours)15.9320.4770.4080.100

Multi-Sequence Generalization

The masked-vs-global advantage generalizes beyond KITTISeq02. Across KITTI 00/027, 02/034, and 05/018, the low-error mask (τ=0.18) outperforms global supervision on every metric, with PSNR gains of +0.44, +0.64, and +0.70 dB at tied or better RMSE.

SeqMethodPSNR ↑SSIM ↑LPIPS ↓RMSE ↓
02/034RGB-only14.900.4330.4460.542
Global (τ=1.0)15.490.4480.4340.101
Masked (τ=0.18)15.930.4770.4080.100
05/018RGB-only14.890.5210.4930.807
Global (τ=1.0)15.260.5340.4690.096
Masked (τ=0.18)15.900.5480.4460.096
00/027RGB-only14.670.4870.4010.514
Global (τ=1.0)16.690.5540.3020.125
Masked (τ=0.18)17.390.5710.2880.114

Comparison to Depth-Supervised Baselines

We benchmark against four depth-supervised methods on KITTISeq02 Frag. 034, all retrained for 50k iterations on our sparse split. The top group uses no GT depth; the bottom group uses GT LiDAR. Our masked DA-V2 approach outperforms three LiDAR-supervised baselines — DNGaussian, DepthRegGS, and SparseGS — which target indoor or object-centric scenes. Only DN-Splatter leads (+0.29 dB PSNR, −0.12 LPIPS), using GT LiDAR unavailable at test time; its margin reflects depth-source quality rather than masking.

MethodPSNR ↑SSIM ↑LPIPS ↓
RGB-only (no depth)14.9030.4330.446
DA-V2 depth, no mask (global)15.4940.4480.434
Ours (τ=0.18, λ=0.10)15.9320.4770.408
DNGaussian (GT LiDAR)9.980.3030.710
DepthRegGS (GT LiDAR)8.710.2290.737
SparseGS (GT LiDAR)12.200.3590.648
DN-Splatter (GT LiDAR)16.220.4890.289

Comparison to a Depth-Inconsistency Mask

We compare our photometric mask against a static, one-shot approximation of the Depth-Inconsistency Mask (DIM) at a matched reliable-pixel fraction. Ratio = MAE(outside) / MAE(inside): >1 means the mask isolates accurate DA-V2 depth; <1 means no separation. Our mask separates accurate from inaccurate depth (ratio 1.17) while the DIM proxy does not (0.87). This is expected: DIM keys on distortions that depth-supervised training itself induces, so a one-shot proxy has nothing to key on. Both are computed once and held fixed, so the comparison is apples-to-apples.

MaskThresholdFractionMAE in / out (m)Ratio
Photometric (ours)τ=0.1895.9%4.14 / 4.821.17
DIM proxyε=17.6 m96.1%4.19 / 3.660.87
Splatfacto lambda by tau qualitative ablation grid
Effect of depth-loss weight λ and reliability threshold τ on Splatfacto. Depth supervision improves geometric consistency and the reconstruction of thin structures, with fewer smeared artifacts around object boundaries and foreground vehicles.

Mip-NeRF-360 — KITTISeq02 every-2

The implicit density field is far more sensitive to noisy monocular depth. The best PSNR (20.607 dB) comes from global (τ=1.0) supervision — the optimal configuration bypasses the mask entirely. The best masked setting trails global by 0.026 dB, well within multi-seed variance (20.57±0.04 dB). All depth-supervised settings, including global, degrade SSIM, LPIPS, and geometry relative to RGB-only (RMSE rises from 2.703 to ≥3.532).

Setting (τ, λ)PSNR ↑SSIM ↑LPIPS ↓RMSE ↓
RGB-only20.3780.6010.4092.703
0.16, 0.1520.5810.5960.4153.654
0.18, 0.1020.3840.5940.4163.532
1.00, 0.15 (global)20.6070.5950.4123.580

Mip-NeRF-360 — KITTISeq05

On KITTISeq05, both global and masked DA-V2 supervision degrade rendering and geometry compared with RGB-only training. Masking recovers 0.22 dB PSNR over global (16.612 vs 16.389), so it mitigates rather than reverses the harm — monocular depth priors are not reliably beneficial for the implicit density representation.

Setting (τ, λ)PSNR ↑SSIM ↑LPIPS ↓AbsRel ↓RMSE ↓
RGB-only17.0680.5460.5290.11662.978
1.00, 0.15 (global)16.3890.5270.5690.15274.803
0.18, 0.1516.6120.5300.5630.13994.454
Mip-NeRF-360 lambda by tau qualitative ablation grid
Effect of depth-loss weight λ and reliability threshold τ on Mip-NeRF-360 reconstruction of finer details such as the street pole. Masked depth supervision provides only weak, unstable regularization for the implicit density field.
PSNR versus depth-loss weight lambda
PSNR as a function of λ on KITTISeq02 across six thresholds τ. Splatfacto peaks sharply at τ=0.18, λ=0.10 (15.93 dB); Mip-NeRF-360 rises only gently with λ.

Object-Centric Bicycle

When multi-view coverage is already strong, the story flips. On the Mip-NeRF-360 Bicycle scene, RGB-only Splatfacto gives the best rendering quality (17.731 PSNR), while depth supervision consistently lowers RMSE (1.479 → 0.722) at the cost of PSNR / SSIM / LPIPS — depth regularization over-constrains an already well-posed reconstruction. Global supervision gives the best PSNR among depth-supervised runs (17.593 vs 17.466 for the best mask), confirming the mask's rendering gain is specific to sparse forward-facing trajectories.

Setting (τ, λ)PSNR ↑SSIM ↑LPIPS ↓RMSE ↓
RGB-only17.7310.5700.2401.479
0.14, 0.1517.0360.4770.3090.722
0.18, 0.0517.4660.5200.2790.751
0.22, 0.1017.1060.4910.2970.725
1.00, 0.05 (global)17.5930.5200.2750.763
Bicycle lambda by tau qualitative ablation grid
Ablation of λ and τ for object-centric Splatfacto reconstruction. Depth supervision improves RMSE but generally reduces RGB rendering quality when multi-view coverage is already strong.

Takeaways

Monocular depth priors are most useful for explicit Gaussian representations in under-constrained, forward-facing sparse-view scenes, and less reliable for implicit NeRF-style density fields. The RMSE drop for Splatfacto comes from using a depth prior at all; masking is what improves rendering fidelity, adding +0.44 to +0.70 dB PSNR over global supervision across KITTI 00/02/05 at tied or better RMSE, with no gain for Mip-NeRF-360 or the object-centric Bicycle scene. Without any ground-truth depth, our mask also outperforms three LiDAR-supervised baselines (DNGaussian, DepthRegGS, SparseGS). Our numbers use DA-V2, but the explicit-vs-implicit architectural insights are prior-agnostic; sweep values (τ, λ) are dataset-specific. Future work: better scale alignment, uncertainty-aware masks, and adaptive depth-loss weighting.

BibTeX

@misc{chu2026reliability,
    title         = {Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction},
    author        = {Wei-Teng Chu and Yashasvini Gopalan and Changju Yuan},
    year          = {2026},
    eprint        = {2607.02554},
    archivePrefix = {arXiv},
    primaryClass  = {cs.CV},
    url           = {https://arxiv.org/abs/2607.02554}
}

Acknowledgements

This project builds on Depth Anything V2, Mip-NeRF 360, Nerfstudio / Splatfacto, and 3D Gaussian Splatting.