PointDiT: Pixel-Space Diffusion for Monocular Geometry Estimation

Haofei Xu
Rundi Wu
Philipp Henzler
Nikolai Kalischek
Michael Oechsle
Fabian Manhardt
Marc Pollefeys
Andreas Geiger
Michael Niemeyer
International Conference on Machine Learning (ICML) (2026)

Abstract

Recent advances in single-image 3D reconstruction have achieved impressive fidelity. However, state-of-the-art methods often retain significant architectural overhead, relying on hybrid backbones or compressed latent representations to ensure performance. In this work, we investigate whether such complexity is strictly necessary. In particular, we introduce a pixel-space diffusion model for single-image point map prediction built upon a plain Vision Transformer (ViT). Unlike leading approaches that depend on specialized hybrid architectures or intricate loss functions, we focus on a minimalist design. We conduct a systematic investigation into the behaviors of simple architectures, identifying the key ingredients required to achieve high-fidelity dense geometry without engineering heavy-handed priors. Through controlled experiments, we demonstrate that this streamlined approach yields results superior to latent-based methods while remaining significantly simpler than hybrid alternatives. By establishing a generic, plain ViT baseline, our framework not only offers a cleaner foundation for 3D prediction but also capitalizes on the scalability of standard vision architectures, allowing it to directly benefit from future architectural advancements and extending its potential to a wider range of 3D geometry tasks.
×