MonoPCA

Rethinking Geometric Depth in
Monocular 3D Object Detection

A Projection-Consistent Reformulation

Zhihao Zhang1 Abhinav Kumar1 Huaizhi Qu2 Tianlong Chen2 Xiaoming Liu1,2

1Michigan State University 2University of North Carolina at Chapel Hill

Papersoon arXivsoon Code BibTeX
NeurIPS 2026
Depth error of GeoDepth vs PCDepth, and the geometric origin of the bias
(a) Even with ground-truth f, H, and h2D, the classic GeoDepth formula incurs a depth error that grows with object size. PCDepth removes most of it. (b) Under perspective projection, the 2D box height encodes the visible front surface, not the object center.
TL;DR

The ubiquitous geometric depth z = fΒ·H / h2D measures the depth of an object's front surface, not its center. We derive a closed-form, parameter-free correction (PCDepth) and pair it with a distillation framework (GeoAlign) that sharpens the size and yaw predictions the correction relies on.

+1.44
AP3D Mod. on the
KITTI leaderboard
+2.80
AP on
Omni3DOUT
+2.32
AP3D Mod. on
nuScenes val
βˆ’13%
GFLOPs vs. MonoCoP
(71.8 β†’ 62.5)

Key Insight

For an object with thickness, the tight 2D box height h2D is set by the box corner nearest to the camera. Plugging it into fΒ·H / h2D therefore returns that corner's depth, z βˆ’ Ξ”, where Ξ” is the box's half-extent along the optical axis. Drag the controls to see how the bias depends on size and yaw.

GeoDepth: front surface PCDepth: object center
GeoDepthβ€”
True centerβ€”
Underestimationβ€”

Idealized geometry (pinhole camera, upright box, tight 2D box), as in Eq. (PCDepth) of the paper. Sizes for the presets are typical values, not dataset statistics.

Abstract

Monocular 3D detection (Mono3D) hinges on recovering depth from a single image, for which the geometric formulation z β‰ˆ fΒ·H / h2D is widely adopted due to its simplicity and direct grounding in projective geometry. We show that this formulation harbors a systematic, projection-induced bias: substituting the 2D bounding-box height h2D recovers the depth of an object's visible front surface rather than its center, causing size-dependent underestimation that persists even under ground-truth inputs. We address this front-surface bias with MonoPCA, a unified framework built on a projection-consistent reformulation. At its core, PCDepth retains the measurable h2D and closes the gap to center depth through an orientation- and size-aware correction derived from perspective geometry, eliminating the bias without introducing ill-posed regression targets. Because PCDepth couples depth to predicted object size and orientation, attribute perception becomes essential. We therefore introduce GeoAlign, a teacher–student framework that transfers geometric representations from Vision Foundation Models into a lightweight student, strengthening attribute perception at no inference cost. Across four challenging benchmarks, MonoPCA achieves state-of-the-art performance, including a +1.44% AP3D gain on the KITTI leaderboard and +2.80% on Omni3DOUT.

Method

Overview of the MonoPCA framework
Overview of MonoPCA. PCDepth corrects the projection bias analytically, making depth an explicit function of the predicted 3D size and yaw. GeoAlign strengthens those attributes by distilling a Mixture-of-Backbone (MoB) teacher into a ResNet student via Hierarchical Feature Alignment (HFA). The teacher is discarded after training.
01

PCDepth

parameter-free

Regressing the center-aligned height hβ€²2D is ill-posed: it depends on the very depth we want. Instead, PCDepth keeps the observable h2D and adds back the optical-axis half-extent of the box footprint in closed form:

Ξ” reduces to W/2 at Ξ³ = 0Β° and L/2 at Ξ³ = 90Β°. Because depth is now computed from predicted attributes, the dedicated depth encoder and depth-guided decoder are no longer needed.

02

GeoAlign

no inference cost

PCDepth ties depth accuracy to size and yaw. GeoAlign builds a high-capacity teacher with a Mixture-of-Backbone that fuses a ResNet pyramid with a Vision Foundation Model (EVA-CLIP) through scale-wise matching and attention, then distills it into a standard ResNet-50 student.

Cosine alignment is invariant to the magnitude gap between a VFM-enriched teacher and a ResNet student, and spans all four pyramid scales.

Geometric basis of PCDepth
Geometric basis of PCDepth. (a) Source of the front-surface bias. (b) hβ€²2D is depth-dependent, so regressing it is ill-posed. (c) Recovering center depth via an optical-axis correction.
Mixture-of-Backbone details
Mixture-of-Backbone. MoB adaptively fuses a convolutional pyramid with a single-scale VFM branch at every scale.

Results

To our knowledge, MonoPCA is the first monocular detector to reach state of the art on both the KITTI leaderboard and Omni3DOUT.

KITTI leaderboard (test) and val, Car, IoU3D β‰₯ 0.7. Bold: best among methods without extra data; underline: second best.

MethodExtra
data
Test AP3DTest APBEVVal AP3DVal APBEV
EasyMod.HardEasyMod.HardEasyMod.HardEasyMod.Hard
OccupancyM3DLiDAR25.5517.0214.7935.3824.1821.3726.8719.9617.1535.7226.6023.68
OPA-3DDepth24.6817.1714.1432.5023.1420.3024.9719.4016.5933.8025.5122.13
MonoTAKDLiDAR27.9119.4316.5138.7527.7624.1434.3622.6119.8842.8629.4126.47
MonoConβ€”22.5016.4613.9531.1222.1019.0026.3319.0115.98β€”β€”β€”
MonoDETRβ€”25.0016.4713.5833.6022.1118.6028.8420.6116.3837.8626.9522.80
MonoUNIβ€”24.7516.7313.49β€”β€”β€”24.5117.1814.01β€”β€”β€”
FD3Dβ€”25.3817.1214.5034.2023.7220.7628.2220.2317.0436.9826.7723.16
MonoMAEβ€”25.6018.8416.7834.1524.9321.7630.2920.9017.6140.2627.0823.14
MonoCDβ€”25.5316.5914.5333.4122.8119.5726.4519.3716.3834.6024.9621.51
MonoDGPβ€”26.3518.7215.9735.2425.2322.0230.7622.3419.0239.4028.2024.42
MonoCoPβ€”27.5419.1116.3336.7725.5722.6232.0623.9820.6441.9230.7527.06
MonoPCA (ours)β€”30.3420.5517.8438.6526.6324.4335.8626.2223.5846.2933.8930.05
Ξ” vs. second best+2.80+1.44+1.06+1.88+1.06+1.81+3.80+2.24+2.94+4.37+3.14+2.99

nuScenes val (front camera), Car. Bold: best; underline: second best.

MethodIoU3D β‰₯ 0.7IoU3D β‰₯ 0.5
AP3DAPBEVAP3DAPBEV
EasyMod.EasyMod.EasyMod.EasyMod.
GUP Net8.507.4014.2112.8129.0326.1633.4230.23
DEVIANT9.698.3316.2814.3631.4728.2235.6131.93
MonoDETR9.538.1916.3914.4131.8128.3535.7031.96
MonoDGP10.048.7816.5514.5329.5626.1732.6729.44
MonoCoP10.859.7117.8315.8633.7029.9137.4434.01
MonoPCA (ours)13.7512.0320.9718.4938.5034.1142.0338.49
Ξ” vs. second best+2.90+2.32+3.14+2.63+4.80+4.20+4.59+4.48

Omni3D (AP averaged over IoU3D thresholds 0.05–0.50) alongside the KITTI leaderboard AP3D.

MethodOmni3DOUTKITTI leaderboard
APkitAPnusAPoutAPlargeAPsmallAPcarEasyMod.Hard
MonoDGPβ€”β€”β€”β€”β€”β€”26.3518.7215.97
MonoCoPβ€”β€”β€”β€”β€”β€”27.5419.1116.33
ImVoxelNet23.523.421.5β€”β€”β€”17.1510.979.15
SMOKE25.920.420.0β€”β€”β€”14.039.767.84
Cube R-CNN36.032.731.937.621.872.923.5915.0112.56
DetAny3D35.833.932.2β€”β€”β€”26.8918.6715.48
MonoPCA (ours)38.137.035.038.825.477.030.3420.5517.84
Ξ” vs. second best+2.1+3.1+2.8+1.2+3.6+4.1+2.80+1.44+1.51

KITTI leaderboard AP3D Mod. vs. model cost. PCDepth is parameter-free and the GeoAlign teacher is discarded after training.

MonoDGP18.7238.9M Β· 69.0 GFLOPs
MonoCoP19.1142.5M Β· 71.8 GFLOPs
MonoPCA20.5542.0M Β· 62.5 GFLOPs
Radar chart comparing MonoPCA with previous state of the art
Across the board. MonoPCA vs. the previous best on each benchmark.
PCDepth improves MonoDETR, MonoDGP and MonoCoP
Plug-and-play. Swapping only the depth head for PCDepth improves MonoDETR, MonoDGP, and MonoCoP by +1.54, +1.37, and +1.09 AP3D on KITTI val.

Which depth formulation?

KITTI val, AP3D Mod. (IoU3D β‰₯ 0.7), full MonoPCA with only the depth formulation changed.

Direct regression23.82
GeoDepth with hβ€²2D24.18
GeoDepth with h2D24.73
PCDepth (ours)26.22

Regressing the unobservable hβ€²2D hurts; keeping the observable h2D and correcting it analytically gives +1.49 over GeoDepth.

Per-distance size, orientation and depth errors
GeoAlign lowers the errors PCDepth depends on. Per-distance size, orientation, and depth errors on KITTI val. Better attributes propagate directly into better depth.

BibTeX

@inproceedings{zhang2026rethinking,
  title     = {Rethinking Geometric Depth in Monocular 3D Object Detection:
               A Projection-Consistent Reformulation},
  author    = {Zhang, Zhihao and Kumar, Abhinav and Qu, Huaizhi and
               Chen, Tianlong and Liu, Xiaoming},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2026}
}