RefineAny3D

Depth Refinement as Semantic Alignment
for Monocular 3D Detection

Zhihao Zhang1 Gengwei Zhang2 Tianlong Chen2 Xiaoming Liu1,2

1Michigan State University 2University of North Carolina at Chapel Hill

Papersoon arXivsoon Code BibTeX
NeurIPS 2026
RefineAny3D overview: depth is the bottleneck, reason-act-refine loop, semantic alignment, and results
(a) Object depth is the bottleneck: ground-truth depth adds +29.92 AP3D on Omni3D, while a depth foundation model (DFM) lowers it by 3.68. (b) A VLM emits action tokens that iteratively adjust depth. (c) Depth error has a visual signature: a too-close box projects too large. (d) Plug-and-play gains on closed-set detection, open-vocabulary detection, and auto-labeling.
TL;DR

Project a candidate 3D box onto the image: if it is too far, the wireframe looks too small; if it is too close, it looks too large. RefineAny3D turns depth refinement into this visual judgment. A vision-language model reasons about the wireframe and emits action tokens (direction + magnitude) instead of a number, and repeats until the box fits.

+3.49
AP3D Mod. on KITTI val
over MonoCoP
+4.35
AP3D on Omni3D
over DetAny3D
+2.37
AP3D Mod. from refined
LabelAny3D labels
1.62
VLM calls per object
on average

Key Insight

Once a 3D box is projected onto the image, camera intrinsics and metric scale are absorbed into the projection. What is left is a purely 2D question: does the wireframe tightly enclose the object? Drag the box depth, then let the refinement loop fix it.

Object (true depth) Candidate box
True depth: 10.0 m ยท object size sobj = 2.40 m
Visual evidence

โ€”

Action tokens

โ€”

    An illustrative simulation of the refinement loop, not the trained model: the action here is computed from the known true depth, and the step sizes are placeholder fractions of sobj. In RefineAny3D, the VLM makes this judgment from the rendered image alone.

    Abstract

    Monocular 3D object detection spans two regimes: closed-set detectors operating within a fixed category vocabulary, and open-vocabulary detectors that localize arbitrary categories by leveraging depth foundation models for 3D geometry. We find that current depth foundation models, despite their strong zero-shot generalization, lack the object-level precision 3D detection demands: substituting a state-of-the-art depth foundation model for a strong detector's predicted depth degrades accuracy. Rather than pushing detectors or depth models to be more accurate end-to-end, we treat object-level depth refinement as a stand-alone task and present RefineAny3D, a vision-language model that corrects depth without ever predicting a numerical value. Our key insight is that depth error has a direct visual signature in image space: when projected onto the image, a correctly placed box tightly encloses the object, while a too-far box projects too small and a too-close box projects too large. Depth refinement thus reduces to a visual alignment problem rather than a metric regression problem, which we instantiate by extending the VLM's vocabulary with action tokens that replace numerical depth output with categorical decisions, and by supervising the model on a large-scale chain-of-thought dataset that grounds each decision in explicit visual evidence. Applied as a single post-hoc step, RefineAny3D delivers consistent gains across closed-set detectors, open-vocabulary detectors, and 3D auto-labeling tools, and generalizes to novel categories, scenes, and cameras without retraining.

    Method

    Overview of the RefineAny3D framework
    Overview of RefineAny3D. The candidate box is rendered as a wireframe on the image. The VLM reasons in four steps (identify, recall, evidence, decision) and emits a direction token and a magnitude token. The depth is updated and the loop repeats until <depth_ok> or a step cap.
    01

    Action tokens

    6 new tokens

    VLMs are unreliable at numerical regression, so RefineAny3D never outputs a number. Its vocabulary is extended with a direction token and a magnitude token, each with its own learnable embedding:

    <depth_closer><depth_ok><depth_farther> <step_small><step_medium><step_large>

    Steps are relative to the object's own size: 0.5 m is negligible for a distant truck but huge for a nearby cup. Precision lost to discretization is recovered by iterating.

    02

    3.0M CoT samples

    from Omni3D

    Training data is curated from Omni3D (six datasets, indoor and outdoor, 50 categories). After geometric and VLM-based filtering, 335K annotations remain. Their depth is perturbed along the camera ray with object-relative, action-balanced magnitudes, giving about 3.0M samples.

    Each sample carries a chain-of-thought rationale before the action, so every decision is grounded in visual evidence rather than a shortcut. Training has two stages: warm up the six token embeddings with the VLM frozen, then fine-tune the LLM with the vision encoder frozen.

    Examples of chain-of-thought supervision
    Chain-of-thought supervision. Each training sample pairs a rendered wireframe with a reasoning trace (identify โ†’ recall โ†’ evidence) that ends in the action tokens.

    Results

    Applied post hoc, with every other box attribute held fixed, RefineAny3D improves closed-set detectors, open-vocabulary detectors, and 3D auto-labeling.

    KITTI val, Car AP3D at IoU3D โ‰ฅ 0.7. RefineAny3D refines the depth of MonoCoP's predictions.

    MethodEasyMod.Hard
    MonoDETR28.8420.6116.38
    MonoCD26.4519.3716.38
    FD3D28.2220.2317.04
    MonoMAE30.2920.9017.61
    MonoTAKD34.3622.6119.88
    MonoDGP30.7622.3419.02
    MonoCoP32.0623.9820.64
    + RefineAny3D (ours)35.6227.4721.31
    ฮ” vs. MonoCoP+3.56+3.49+0.67

    Omni3D test, AP3D per sub-dataset. RefineAny3D refines DetAny3D conditioned on ground-truth 2D boxes, its strongest variant; only depth changes.

    MethodKITTInuScenesSUN RGB-DARKitObjectronHypersimOverall
    Cube R-CNN32.5030.0615.3341.7350.847.4823.26
    OVMono3D25.4524.3315.2041.6058.877.7522.98
    DetAny3D31.6130.9718.9646.1354.427.1724.92
    DetAny3D w/ GT 2D box38.6837.5546.1450.6256.8215.9834.38
    + RefineAny3D (ours)43.4744.1348.5854.2059.2917.8038.73
    ฮ” vs. DetAny3D w/ GT 2D+4.79+6.58+2.44+3.58+2.47+1.82+4.35

    KITTI val AP3D of a detector (MonoCoP) trained on each set of labels. Refining LabelAny3D's pseudo-labels narrows the gap to ground truth.

    Training labelsEasyMod.Hard
    Ground truth (upper bound)32.0623.9820.64
    LabelAny3D13.5711.7810.27
    + RefineAny3D (ours)15.7614.1510.99
    ฮ” vs. LabelAny3D+2.19+2.37+0.72

    Refine3D, our standalone benchmark on Omni3D. DirAcc: direction correct; FullAcc: direction and magnitude correct; DepthErr: residual depth gap (lower is better).

    MethodStandardNovel categoryNovel camera
    DirAccFullAccDepthErrDirAccFullAccDepthErrDirAccFullAccDepthErr
    Qwen3-VL-8B28.711.40.3227.510.90.3728.311.10.33
    Qwen3-VL-8B w/ our data75.655.40.2156.341.20.2768.749.80.23
    RefineAny3D (ours)89.276.70.1674.458.50.1985.673.20.17

    Refining MonoCoP's depth on KITTI val (AP3D, IoU3D โ‰ฅ 0.7). Geometric fitting even uses the ground-truth 2D box.

    MethodEasyMod.Hard
    MonoCoP32.0623.9820.64
    Geometric fitting29.8921.5518.36
    Geometric fitting (GT dims + yaw)30.5621.7918.43
    DINOv2 action classifier24.7117.6212.78
    RefineAny3D (ours)35.6227.4721.31

    Both simpler judges make MonoCoP worse. The gain comes from the VLM's visual judgment, not the discrete action space alone.

    AP3D versus number of refinement iterations on KITTI Moderate
    Fast convergence. On KITTI Moderate, one step recovers about 85% of the gain (+2.95); a second adds +0.54, after which accuracy plateaus. We cap refinement at two steps.

    What matters in the design?

    Refine3D standard split, FullAcc (direction and magnitude both correct). Each row changes one choice.

    Plain-text actions55.4
    One joint token per action72.5
    No chain-of-thought70.2
    Unfrozen vision encoder69.6
    RefineAny3D (dir + mag)76.7

    Factored direction and magnitude tokens each see about a third of the training samples, versus about a seventh for joint tokens. Chain-of-thought and a frozen vision encoder both matter too.

    BibTeX

    @inproceedings{zhang2026refineany3d,
      title     = {RefineAny3D: Depth Refinement as Semantic Alignment
                   for Monocular 3D Detection},
      author    = {Zhang, Zhihao and Zhang, Gengwei and Chen, Tianlong and
                   Liu, Xiaoming},
      booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
      year      = {2026}
    }