Lucida AI
Research preview - Lucida real-to-sim scene modeling

Turn a real room into an editable 3D scene

Lucida AI takes a cluttered room capture and reconstructs it as a scene graph of complete, editable 3D assets - so robots and embodied agents can train in a digital twin of the real world.

Free to start - No credit card required - Lucida AI powered

Lucida AI real-to-sim 3D scene workspace

One video in. A scene you can edit.

Lucida splits real-to-sim into three dedicated stages - parse, generate and place - so every object comes back as a complete, independent and correctly positioned asset.

Video-based capture

Walk a camera around a room and Lucida builds a scene-level understanding of every object in view.

Scene graph parsing

Each instance is anchored with per-instance multi-view evidence instead of a flat object list.

Complete editable assets

Objects come back fully modeled, including the parts the camera could not see.

Closed-loop placement

A VLM agent places each asset and self-checks alignment until the scene is consistent.

Simulation ready

Export scenes for robot manipulation, navigation and embodied AI training loops.

Benchmark leading

69% higher scene-level detection mAP over prior methods on R2S-Scene, with 0.924 scene F-Score.

Understand the room as a scene graph

Lucida parses the input video into a structured scene graph where every node carries the visual evidence gathered from multiple viewpoints - categories, poses and spatial relations included.

Per-instance evidence

Multi-view observations make each object reliable to detect and reconstruct.

Relationships preserved

Spatial context stays intact, so a lamp is not just a lamp - it is a lamp on the side table.

Works from clutter

Real captures with occlusions and messy layouts are handled by design.

Editable assets, not dead meshes

For every detected instance, Lucida generates a complete, editable asset - editable 3DGS or mesh - so hidden geometry, missing textures and unfinished backs are no longer a blocker.

Occlusion aware

Geometry the capture never saw is inferred to complete the object.

Editable output

Each object stays an independent asset you can scale, retarget or replace.

Standard formats

Export Ed-3DGS or mesh representations that drop straight into your simulator.

GizmoAct places every object in its place

Lucida models placement as multi-round GUI interaction: a VLM agent manipulates the gizmo of each asset and closes the loop by judging its own alignment against the scene.

Interacts like a human

Move, rotate and scale through familiar controls instead of one-shot regressions.

Self-correcting

Alignment is checked visually and re-adjusted until it is right.

Physically grounded

Objects land on surfaces with plausible supports - ready for a robot to interact with.

Results that make simulation trustworthy

Evaluated on R2S-Scene, CA-1M and real-to-sim scene reconstruction benchmarks.

69% higher scene-level mAP over Boxer on R2S-Scene

69%

higher scene-level mAP over Boxer on R2S-Scene

83.4% ADD-SB at 0.05 threshold on CA-1M object poses

83.4%

ADD-SB at 0.05 threshold on CA-1M object poses

0.924 scene F-Score, up from 0.794 with SAM3D

0.924

scene F-Score, up from 0.794 with SAM3D

3 dedicated stages: parse, generate and place

3

dedicated stages: parse, generate and place

FAQ

Everything you need to know about Lucida AI.







Turn your next capture into a sim-ready scene

Lucida AI is in research preview. Get early access and start building editable digital twins today.

Lucida AI