A human-native multimodal world model for embodied AI.
Let robots learn how actions change the world — by watching people go about their day.
01 / PRODUCT
Turning everyday human activity into robot-trainable data.
- MIRO-EGO01

Head-mounted capture kit
A head unit and two sEMG armbands record stereo RGB, depth, IMU, gaze, speech and muscle activation in sync. 200 grams and about USD 280 per unit — no capture cell, no optical mocap, one wearer in a real scene is a data production line.
- MIRO-CAP02

Multi-view capture and reconstruction
An egocentric view plus several static cameras, no mocap studio. Video goes in; people, hands, objects and contact come out in one world frame, then retarget onto a humanoid, a dexterous hand or a gripper and run in simulation.
- MIRO-HAND03

Egocentric hand reconstruction model
One model for a single frame and for variable-length video. RGB is the main input, with depth and camera parameters added when available; given the camera pose, it outputs both hands as MANO in world space.
Reconstruction accuracy
Hand world-frame error 22.9mm
Measured on public HOI4D, 47% below the runner-up.
Object rotation error 12.18°
23% below the runner-up on the same benchmark. Baselines use their own published results.
Simulation tracking success 99.9%
A humanoid trained on the reconstructed data, captured with two iPhones.
02 / ABOUT
Our core team brings years of experience in 3D human modeling and embodied AI.
- 20,000+Google Scholar citations, combined
- 10,000+GitHub stars across open-source work
- 60Mframes of 4D human data
- Top CS venuesBest Paper and Test-of-Time awards
03 / CONTACT
Contact us
Whether you're building robots or doing related research, we'd love to talk.
We are hiring in 3D vision, machine learning and embodied AI — Join us