Action tokens
6 new tokensVLMs are unreliable at numerical regression, so RefineAny3D never outputs a number. Its vocabulary is extended with a direction token and a magnitude token, each with its own learnable embedding:
Steps are relative to the object's own size: 0.5 m is negligible for a distant truck but huge for a nearby cup. Precision lost to discretization is recovered by iterating.


