Single-image 3D scene reconstruction

Building Rome from a Single Image

1Applied Intuition 2Purdue University 3UIUC 4UC Berkeley

Paper Code (soon)

Interactive

Explore the reconstructions

Drag to orbit, scroll to zoom, WASD to move, QE down/up. Measure: pick two points to get their distance in metres.

Note: meshes are simplified to about 1.5M faces for this web viewer; the paper figures and videos use the full-resolution meshes.

Overview

An object generator, redesigned for whole scenes

  • Adaptive scene chunks. Chunks grow with distance: small ones keep nearby detail, large ones cover distant buildings.
  • Explicit 2D–3D conditioning. Image features are lifted onto the observed surface, and each voxel knows if it is free, visible, or hidden.
  • Autoregressive generation. Neighbouring chunks reuse overlapping latents, so new geometry stays consistent with the scene.
Pipeline: (a) the image is lifted with a point map and split into adaptive 3D chunks; (b) each chunk is generated by a sparse-structure transformer conditioned on lifted features and visibility, then a structured-latent transformer; (c) chunks are generated autoregressively with reused latents and decoded into one scene mesh.

Comparisons

Input photo, scene 1
Input
Ours
World Tracing
GenRecon
VolFill
Lyra 2.0