Rigid · Articulated · Deformable — sim-ready scenes from multi-view RGB

CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes with Deformable Curves, Surfaces, and Volumes

  • 1SIGS, Tsinghua University
  • 2Carnegie Mellon University
  • 3Genesis AI

*Equal contribution.  †Corresponding authors.

1DCurve · telephone cord → rod
2DSurface · plastic bag → shell
3DVolume · chair cushion → solid

Overview

Teaser: input views of an office (left), the reconstructed sim-ready scene with a chair cushion, telephone cord, plastic bag in a rigid bin, and a sliding drawer (center), and robot interactions with the curve, surface, and volume (right).

CoDimRecon is an agentic framework that turns multi-view RGB images into an editable, simulation-ready scene of rigid, articulated, and deformable objects. Scene-level geometric priors ground scale and layout, object-level generated meshes guide the agent toward detailed, compact geometry, and articulated objects such as the office drawer above are split into movable parts with explicit joints.

Deformables are reconstructed by dimensionality in dedicated agent sessions: curves such as the telephone cord become centerlines with radii, surfaces such as the plastic bag become manifold shells with thickness — kept separate from its rigid bin — and volumes such as the chair cushion become watertight solids for volumetric meshing. Reusable simulator skills then initialize compatible physical models and parameters, and agent-guided robot tests expose mismatches and trigger targeted revisions of motion, geometry, numerics, or material modeling, yielding a scene that robots can directly interact with in simulation.

Simulation-ready assets · Scene 067-D

Inspect the physical properties

Every object in the reconstructed office is handed to the simulator with a physical description. Click any object to see its category, mass, friction, and joints, and for deformable curves, surfaces, and volumes, the constitutive model and parameters used in simulation.

Preparing the reconstructed scene
Loads when you scroll here
Highlight

Deformable simulation

None of the evaluated reconstruction baselines outputs deformable assets for physical simulation, so we report capability demonstrations. Reconstructed curves, surfaces, and volumes are simulated with an inserted robot on an IPC-family contact backend and replayed in Blender for rendering.

Curve · discrete elastic rod

Telephone cord

A Franka lifts the handset through finger contact and the coiled cord extends. The rod keeps its reconstructed coil as natural shape; peak handset lift is about 198 mm.

Surface · thin shell

Plastic bag in a rigid bin

The 0.03 mm bag deforms freely with no pinned vertices and is held through contact alone; the bin is a separate rigid body.

Surface · plastic hinges

Robotic paper folding

The gripper folds the sheet through 140°. With plastic hinge bending selected by behavioral testing, the crease is retained after release.

Volume · StVK–Hencky solid

Chair cushion press

A 70 mm rounded tool presses the cushion at three points; it indents by about 23 mm and recovers. Material and lighting are modified to make the indentation visible.

Volume · stable Neo-Hookean

Beanbag rest shape

Released under gravity without supports or damping, the reconstructed beanbag settles into its equilibrium rest shape.

Articulated · prismatic joint

Drawer sliding

The robot pulls the drawer open along the sliding direction recovered during articulated-object reconstruction.

Embodied · humanoid

Humanoid walking

A G1 humanoid walks through the reconstructed office, which serves directly as a simulation environment.

Method

CoDimRecon takes multi-view RGB observations and proceeds in three stages. The agent first authors an editable scene using geometric context and generated meshes as references, then refines appearance, articulation, and rigid-body stability. Finally, it reconstructs deformable curves, surfaces, and volumes and tests their behavior through simulated robot interaction.

Method overview in three panels: (a) scene reconstruction from multi-view RGB with 3D and semantic priors and mesh references rebuilt as editable primitives; (b) scene refinement with a render–evaluate–refine loop and articulated objects such as a sliding drawer and pressable keyboard; (c) category-wise processing of curves, surfaces, and volumes and an agent-guided behavioral verification loop.
Overview. (a) Geometric context and generated meshes guide editable primitive-based scene reconstruction. (b) Render–evaluate–refine improves appearance and pose; articulated rigid objects receive joints, and rigid bodies are settled under gravity. (c) Separate sessions reconstruct solver-compatible curves, surfaces, and volumes; reusable skills initialize physical models, and agent-guided robot tests diagnose motion, geometry, or numerical issues before material changes.
a

Reconstruction with geometric & generative references

  • Geometric context. VGGT-Omega provides intrinsics, camera poses, depth, and point maps that the agent queries to ground object dimensions, poses, and layout.
  • Mesh references. CropFormer masks are clustered across views into 3D tracks; SAM3D generates a mesh per track, registered to the scene. Small missing objects are recovered with REST3D.
  • Primitive fitting. In Blender, the agent rebuilds furniture from primitives with bevel, subdivision, and lattice modifiers — compact, editable, and ready for articulation.
b

Scene refinement

  • Render–evaluate–refine. A quadratic color alignment c′ = a c² + b c + d acts as a diagnostic: large gains point to lighting or materials, small gains to geometry or pose.
  • Articulation. A joint-modeling guide covers revolute, prismatic, screw, cylindrical, universal, and spherical joints — down to keyboard keys and telephone buttons.
  • Rigid-body stabilization. The scene is settled under gravity in MuJoCo and the resulting poses are written back as canonical placements.
c

Category-wise deformable reconstruction & simulation

  • Separate agent sessions for curves, surfaces, and volumes; combining all instructions degrades curve reconstruction.
  • Reusable simulator skills package constitutive models, runnable example scenes, scripts, and candidate parameters on an IPC-family contact backend.
  • Behavioral verification. Preset robot tests check physical validity and task-level acceptance; failures trigger targeted revisions.

Representation follows dimensionality

Curves, surfaces, and volumes — each in its solver-native form

Cables, cloth, and paper are thin: resolving their thickness volumetrically needs fine through-thickness resolution and can suffer locking. Codimensional models instead represent rods as curves and shells as surfaces embedded in 3D — which changes the reconstruction target itself.

Curves 1D · rod

e.g. the coiled telephone cord, cables

Geometry
Continuous centerline with radius
Repair
Path completion and radius estimation
Simulation
Discrete elastic rod with stretching, bending, and twist; the reconstructed coil is the stress-free natural shape

Surfaces 2D · shell

e.g. paper, plastic bags, clothing

Geometry
Manifold shell (midsurface) with thickness
Repair
Hole patching and fragment merging
Simulation
StVK membrane with hinge bending; plastic hinges when the task needs a lasting crease

Volumes 3D · solid

e.g. chair cushions, beanbags

Geometry
Watertight solid for tetrahedral meshing
Repair
Watertight closure; high-poly assets are bound to a closed low-poly proxy
Simulation
Solid finite elements — StVK–Hencky for large compression, stable Neo-Hookean

Closing the reconstruction-to-simulation loop

Agent-guided behavioral verification

Simulator examples only initialize a deformable asset. Whether it supports the intended interaction becomes evident only when it is exercised, so the agent plans diagnostic robot manipulations, reviews each rollout, and revises in a fixed order — changing the material model only as a last resort.

  1. InitializeModel & parameters from reusable skills

    Geometry fixes the discretization; the agent picks a material model and adapts the closest example scene.

  2. Rest shapeSettle under gravity

    Excessive sag sends the geometry back for revision; the equilibrium is written back to Blender.

  3. Plan & simulatePreset diagnostic manipulation

    Extend the cord, press the cushion, fold the paper. The agent sets grasps, approach poses, and arm motion.

  4. CheckValidity + task acceptance

    Bounds on element stretch and inversion, plus task criteria such as lift height or retained fold.

✓ Pass → accept & freeze

Accepted values are effective parameters under the assumed model — not measured material properties.

✗ Fail → diagnose in a fixed order

  1. Motioncommanded motion, tool, contact location
  2. Geometryshape and boundary conditions
  3. Numericssolver and contact settings
  4. Material modelonly if behavior still can't be reproduced

Compositional scene reconstruction

We evaluate three Replica and three ScanNet++ scenes from the HoloScene release against HoloScene (optimization-based), ReplicateAnyScene (zero-shot compositional), and GPT-6 Astra — an agent-only baseline that shares our backbone but authors each scene in Blender directly from RGB frames.

HoloScene's higher PSNR and SSIM come with visibly fragmented geometry. Pixel metrics penalize the lighting and material differences that remain in our explicitly authored scenes, whereas in the feature space of a frozen video encoder our renders are closer to the source than HoloScene's on both datasets — suggesting the gap is mainly appearance rather than missing or misplaced content.

ScanNet++ Replica
Click the figure to enlarge
Qualitative comparison on ScanNet++ scene 67d702f2e8: rendered appearance (top row) and geometry (bottom row) for ground truth, HoloScene, ReplicateAnyScene, GPT-6 Astra, and CoDimRecon.

Positioning

Capabilities at a glance

MethodInputDeformable simulation Geometric guidanceGenerative priorAuto instance discovery Training-freeAgentic
HoloSceneRGB, mask, cam✗✓✓✗✗✗
SimReconRGB✗✓✓✓✓✗
ReplicateAnySceneRGB✗✓✓✓✓✗
VIGASingle-view RGB✗✗✓✓✓✓
LucidaRGB✗✓✓✓✗✓
LumeraSingle-view RGB✗✓✓✓✗✓
LiteReality-AgentPosed RGBD✗✓✓✓✓✓
GPT-6 AstraRGB✗✗✗✓✓✓
CoDimRecon (ours)RGB✓✓✓✓✓✓

Comparison of compositional scene-reconstruction methods. Cam: camera parameters; RGBD: RGB images with depth. None of the compared methods reconstructs deformable curves, surfaces, and volumes for physical simulation.

BibTeX

@article{xie2026codimrecon,
  title   = {CoDimRecon: Agentic Reconstruction of Sim-Ready 3D Scenes
             with Deformable Curves, Surfaces, and Volumes},
  author  = {Xie, Shuzhao and Wang, Lelin and Lin, Guying
             and Wang, Zhi and Li, Minchen},
  journal = {arXiv preprint arXiv:2609.36024},
  year    = {2026}
}

Acknowledgments

We thank Yi-Ling Qiao, Xinyu Lu, and Kemeng Huang for their guidance on using the IPC-based simulator. This work is supported in part by the National Natural Science Foundation of China (Grant Nos. 92467204 and 62472249), the Shenzhen Science and Technology Program (Grant No. KJZD20240903102300001), and gift funding from Genesis AI. Shuzhao Xie's work is supported by the Google Cloud Research Credits program. Shuzhao Xie thanks Chen Tang for help with computational resources.