Sparse-view neural reconstruction is challenging in outdoor driving scenes, where cameras move along a narrow forward-facing trajectory and provide limited multi-view overlap. Monocular depth estimators can provide dense geometric priors, but their predictions are noisy and not uniformly reliable across image regions. We use Depth Anything V2 (DA-V2) as a dense monocular depth prior, align its predictions to metric depth via per-image scale-shift fitting on sparse anchors (LiDAR on KITTI, COLMAP on Bicycle), and apply depth supervision selectively through photometric masks generated from an RGB-only baseline. We evaluate this strategy on two representative scene representations: Mip-NeRF-360 and Splatfacto. Since the mask modifies an existing depth loss rather than introducing one, we compare against global (unmasked) supervision as the primary baseline. Across KITTI sequences 00/02/05, the mask yields +0.44 to +0.70 dB PSNR for Splatfacto over global supervision at tied or better RMSE, and no change on Mip-NeRF-360 — the mask primarily enhances rendering fidelity rather than geometry. Matched-ratio ablations confirm the gains come from selecting reliable low-error regions rather than simply reducing the number of depth-supervised pixels. Without any ground-truth depth, our mask also outperforms three LiDAR-supervised baselines. On the object-centric Mip-NeRF-360 Bicycle scene, depth supervision improves geometry but hurts rendering quality when multi-view coverage is already strong.
DA-V2 predicts relative depth, so we align each prediction to a metric reference before training. For KITTI the reference is valid LiDAR depth; for the Bicycle scene it is sparse COLMAP points. Given a monocular prediction \(d_m(u)\) and reference depth \(d_r(u)\) over valid anchor pixels \(u \in \Omega\), we solve a per-image least-squares scale-shift fit:
The aligned prior is \(\hat{d}(u) = s^{*} d_m(u) + t^{*}\), clipped to the 80 m evaluation range on KITTI.
Rather than supervising every pixel, we only supervise where the RGB-only baseline is already reliable. From a baseline render \(\hat{I}\) and the ground-truth image \(I\), we compute a per-pixel photometric error and threshold it at \(\tau\):
High photometric error flags unreliable regions — dynamic objects, reflective surfaces, occlusion boundaries, and sky — where a depth loss would push the model toward wrong geometry. Low-error pixels are already consistent with the RGB observations, so the aligned prior acts as a stabilizing regularizer. The mask is computed once from the RGB-only baseline and held fixed throughout depth-supervised training.
The effective supervision mask combines the photometric mask with the depth-validity mask \(D(u)\): \(M_\text{eff} = M_\tau \wedge D\). The depth loss is gated by \(M_\text{eff}\), while the RGB loss is evaluated over the full image:
Setting \(\tau=1.0\) gives \(M_\tau \equiv 1\) — global (unmasked) depth supervision. This condition, not the depth-free \(\lambda_\text{depth}=0\) baseline, is the primary baseline for judging the mask's contribution: it isolates the effect of where depth is supervised from the effect of whether depth is supervised at all.
Before using DA-V2 as supervision, we evaluate the aligned depth against valid KITTI LiDAR pixels. The mean absolute error of alignment is 4.22 m. Sweeping the number of LiDAR anchors used for scale-shift fitting shows the 2-DOF optimization saturates long before typical LiDAR density is reached — alignment MAE is stable from ~95k anchors down to 500 per frame, degrading only under extreme sparsity. The 4.22 m error floor reflects DA-V2's intrinsic local structural limitations rather than calibration artifacts.
| Anchors / frame | Alignment MAE ↓ (m) |
|---|---|
| ~95,000 | 4.22 |
| 5,000 | 4.22 |
| 500 | 4.22 |
| 100 | 4.29 |
| 50 | 4.29 |
| 20 | 4.51 |
Because photometric and depth errors are measured against different references, low photometric error does not by construction imply low depth error. We validate the mask directly against GT LiDAR. Depth RMSE is 34–39% lower inside the mask at every threshold, and inside-mask RMSE increases monotonically with τ (r=1.0). Pixel-level differences are highly significant (p<10−50); the frame-level paired t-test gives p=0.047 (n=8).
| τ | Inside RMSE ↓ (m) | Outside RMSE ↓ (m) |
|---|---|---|
| 0.14 | 6.30 | 9.58 |
| 0.16 | 6.32 | 9.82 |
| 0.18 (ours) | 6.33 | 10.01 |
| 0.20 | 6.36 | 10.24 |
| 0.22 | 6.38 | 10.45 |
| 1.00 (global) | 6.43 | — |
Splatfacto gains clearly from monocular depth supervision. Using any DA-V2 depth prior — even without a mask — improves PSNR (14.903 → 15.494 for global) and cuts RMSE by ~80% (0.542 → 0.101). Masking the depth loss at τ=0.18 adds a further +0.44 dB PSNR at tied RMSE. In short: the RMSE drop comes from using a depth prior at all, and masking is what improves rendering fidelity. Over three seeds, masked 15.30±0.57 vs global 15.07±0.38 dB (+0.23 dB mean); the −80% RMSE improvement over RGB-only is consistent across seeds.
| Setting (τ, λ) | PSNR ↑ | SSIM ↑ | LPIPS ↓ | RMSE ↓ |
|---|---|---|---|---|
| RGB-only | 14.903 | 0.433 | 0.446 | 0.542 |
| 0.16, 0.10 | 15.452 | 0.454 | 0.434 | 0.106 |
| 0.18, 0.10 (ours) | 15.932 | 0.477 | 0.408 | 0.100 |
| 0.18, 0.15 | 15.588 | 0.459 | 0.436 | 0.096 |
| 1.00, 0.10 (global) | 15.494 | 0.448 | 0.434 | 0.101 |
To test whether the benefit comes from selecting reliable pixels or simply using fewer pixels, we compare against high-error and random masks with the identical per-frame pixel budget. The low-error mask wins on every metric — the improvement is not from fewer supervised pixels.
| Mask type (λ=0.10) | PSNR ↑ | SSIM ↑ | LPIPS ↓ | RMSE ↓ |
|---|---|---|---|---|
| High-error, matched | 14.932 | 0.437 | 0.455 | 0.111 |
| Random, matched (3 seeds) | 15.036 | 0.442 | 0.456 | 0.109 |
| Low-error, τ=0.18 (ours) | 15.932 | 0.477 | 0.408 | 0.100 |
The masked-vs-global advantage generalizes beyond KITTISeq02. Across KITTI 00/027, 02/034, and 05/018, the low-error mask (τ=0.18) outperforms global supervision on every metric, with PSNR gains of +0.44, +0.64, and +0.70 dB at tied or better RMSE.
| Seq | Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ | RMSE ↓ |
|---|---|---|---|---|---|
| 02/034 | RGB-only | 14.90 | 0.433 | 0.446 | 0.542 |
| Global (τ=1.0) | 15.49 | 0.448 | 0.434 | 0.101 | |
| Masked (τ=0.18) | 15.93 | 0.477 | 0.408 | 0.100 | |
| 05/018 | RGB-only | 14.89 | 0.521 | 0.493 | 0.807 |
| Global (τ=1.0) | 15.26 | 0.534 | 0.469 | 0.096 | |
| Masked (τ=0.18) | 15.90 | 0.548 | 0.446 | 0.096 | |
| 00/027 | RGB-only | 14.67 | 0.487 | 0.401 | 0.514 |
| Global (τ=1.0) | 16.69 | 0.554 | 0.302 | 0.125 | |
| Masked (τ=0.18) | 17.39 | 0.571 | 0.288 | 0.114 |
We benchmark against four depth-supervised methods on KITTISeq02 Frag. 034, all retrained for 50k iterations on our sparse split. The top group uses no GT depth; the bottom group uses GT LiDAR. Our masked DA-V2 approach outperforms three LiDAR-supervised baselines — DNGaussian, DepthRegGS, and SparseGS — which target indoor or object-centric scenes. Only DN-Splatter leads (+0.29 dB PSNR, −0.12 LPIPS), using GT LiDAR unavailable at test time; its margin reflects depth-source quality rather than masking.
| Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|
| RGB-only (no depth) | 14.903 | 0.433 | 0.446 |
| DA-V2 depth, no mask (global) | 15.494 | 0.448 | 0.434 |
| Ours (τ=0.18, λ=0.10) | 15.932 | 0.477 | 0.408 |
| DNGaussian (GT LiDAR) | 9.98 | 0.303 | 0.710 |
| DepthRegGS (GT LiDAR) | 8.71 | 0.229 | 0.737 |
| SparseGS (GT LiDAR) | 12.20 | 0.359 | 0.648 |
| DN-Splatter (GT LiDAR) | 16.22 | 0.489 | 0.289 |
We compare our photometric mask against a static, one-shot approximation of the Depth-Inconsistency Mask (DIM) at a matched reliable-pixel fraction. Ratio = MAE(outside) / MAE(inside): >1 means the mask isolates accurate DA-V2 depth; <1 means no separation. Our mask separates accurate from inaccurate depth (ratio 1.17) while the DIM proxy does not (0.87). This is expected: DIM keys on distortions that depth-supervised training itself induces, so a one-shot proxy has nothing to key on. Both are computed once and held fixed, so the comparison is apples-to-apples.
| Mask | Threshold | Fraction | MAE in / out (m) | Ratio |
|---|---|---|---|---|
| Photometric (ours) | τ=0.18 | 95.9% | 4.14 / 4.82 | 1.17 |
| DIM proxy | ε=17.6 m | 96.1% | 4.19 / 3.66 | 0.87 |
The implicit density field is far more sensitive to noisy monocular depth. The best PSNR (20.607 dB) comes from global (τ=1.0) supervision — the optimal configuration bypasses the mask entirely. The best masked setting trails global by 0.026 dB, well within multi-seed variance (20.57±0.04 dB). All depth-supervised settings, including global, degrade SSIM, LPIPS, and geometry relative to RGB-only (RMSE rises from 2.703 to ≥3.532).
| Setting (τ, λ) | PSNR ↑ | SSIM ↑ | LPIPS ↓ | RMSE ↓ |
|---|---|---|---|---|
| RGB-only | 20.378 | 0.601 | 0.409 | 2.703 |
| 0.16, 0.15 | 20.581 | 0.596 | 0.415 | 3.654 |
| 0.18, 0.10 | 20.384 | 0.594 | 0.416 | 3.532 |
| 1.00, 0.15 (global) | 20.607 | 0.595 | 0.412 | 3.580 |
On KITTISeq05, both global and masked DA-V2 supervision degrade rendering and geometry compared with RGB-only training. Masking recovers 0.22 dB PSNR over global (16.612 vs 16.389), so it mitigates rather than reverses the harm — monocular depth priors are not reliably beneficial for the implicit density representation.
| Setting (τ, λ) | PSNR ↑ | SSIM ↑ | LPIPS ↓ | AbsRel ↓ | RMSE ↓ |
|---|---|---|---|---|---|
| RGB-only | 17.068 | 0.546 | 0.529 | 0.1166 | 2.978 |
| 1.00, 0.15 (global) | 16.389 | 0.527 | 0.569 | 0.1527 | 4.803 |
| 0.18, 0.15 | 16.612 | 0.530 | 0.563 | 0.1399 | 4.454 |
When multi-view coverage is already strong, the story flips. On the Mip-NeRF-360 Bicycle scene, RGB-only Splatfacto gives the best rendering quality (17.731 PSNR), while depth supervision consistently lowers RMSE (1.479 → 0.722) at the cost of PSNR / SSIM / LPIPS — depth regularization over-constrains an already well-posed reconstruction. Global supervision gives the best PSNR among depth-supervised runs (17.593 vs 17.466 for the best mask), confirming the mask's rendering gain is specific to sparse forward-facing trajectories.
| Setting (τ, λ) | PSNR ↑ | SSIM ↑ | LPIPS ↓ | RMSE ↓ |
|---|---|---|---|---|
| RGB-only | 17.731 | 0.570 | 0.240 | 1.479 |
| 0.14, 0.15 | 17.036 | 0.477 | 0.309 | 0.722 |
| 0.18, 0.05 | 17.466 | 0.520 | 0.279 | 0.751 |
| 0.22, 0.10 | 17.106 | 0.491 | 0.297 | 0.725 |
| 1.00, 0.05 (global) | 17.593 | 0.520 | 0.275 | 0.763 |
Monocular depth priors are most useful for explicit Gaussian representations in under-constrained, forward-facing sparse-view scenes, and less reliable for implicit NeRF-style density fields. The RMSE drop for Splatfacto comes from using a depth prior at all; masking is what improves rendering fidelity, adding +0.44 to +0.70 dB PSNR over global supervision across KITTI 00/02/05 at tied or better RMSE, with no gain for Mip-NeRF-360 or the object-centric Bicycle scene. Without any ground-truth depth, our mask also outperforms three LiDAR-supervised baselines (DNGaussian, DepthRegGS, SparseGS). Our numbers use DA-V2, but the explicit-vs-implicit architectural insights are prior-agnostic; sweep values (τ, λ) are dataset-specific. Future work: better scale alignment, uncertainty-aware masks, and adaptive depth-loss weighting.
@misc{chu2026reliability,
title = {Reliability-Aware Monocular Depth Supervision for Sparse-View Neural Reconstruction},
author = {Wei-Teng Chu and Yashasvini Gopalan and Changju Yuan},
year = {2026},
eprint = {2607.02554},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2607.02554}
}
This project builds on Depth Anything V2, Mip-NeRF 360, Nerfstudio / Splatfacto, and 3D Gaussian Splatting.