AikanAI
Back to feed
News feedGlobalrobotics

Interactive world simulator for robot policy training and evaluation

SourceRobohub(robohub.org)Jul 20, 2026 · 7/20/2026
Interactive world simulator for robot policy training and evaluation

Imagine you want to teach a robot to push an object on a table. The standard recipe in robot learning is to collect hundreds of expert demonstrations on a real robot, train an imitation learning policy on that data, and then evaluate the policy by running it many times on the same real robot. Both stages (data collection and evaluation) are slow, expensive, and hard to reproduce: hardware breaks, lighting changes, objects drift out of place, and every new task means more hours in the lab.

A natural question is whether we can replace some of this real-robot work with a simulator. Classical physics-based simulators are powerful, but building one for a new task means manually modeling geometries, contacts, friction, and deformation, and the resulting simulator often still does not match reality closely enough for policies trained inside it to transfer.

In our work, we take a different route. We build an Interactive World Simulator : a learned, action-conditioned video prediction model that, given the current image and a sequence of robot actions, predicts the next frames purely in pixel space, with no physics engine inside. You can plug in a teleoperation device and control the robot through this learned world model for more than 10 minutes at 15 FPS on a single RTX 4090, and the predicted video stays stable and physically plausible.

https://aihub.org/wp-content/uploads/2026/07/twitter_1.mp4

The key idea is that, if the simulator is faithful enough, we could unlock two long-standing bottlenecks in robot learning:

– Data generation for training becomes cheap, because we can collect demonstrations inside the simulator.

– Policy evaluation becomes scalable and reproducible, because we can roll many policies through the simulator under identical conditions.

https://aihub.org/wp-content/uploads/2026/07/twitter_2.mp4

What the world simulator can do

We trained our world simulator on four manipulation tasks that span very different physical regimes: T pushing (rigid-body contact), rope routing (deformable–rigid interaction with a clip), mug grasping (fine-grained gripper dynamics), and pile sweeping (manipulating piles of objects). All four behaviors are learned from interaction data alone, with no physics priors hard-coded.

A few examples of what the model captures:

Rope routing: it correctly distinguishes between the rope actually being inserted into the clip and the rope swinging past it without making contact. Crucially, it does not bias toward either outcome — it follows what the actions imply.

https://aihub.org/wp-content/uploads/2026/07/twitter_5.mp4

Mug grasping: it captures fine-grained effects such as the mug slipping out of the gripper, or the handle being nudged and rotated.

https://aihub.org/wp-content/uploads/2026/07/twitter_6.mp4

Pile sweeping: it can generate video for multiple viewpoints consistently

https://aihub.org/wp-content/uploads/2026/07/twitter_8.mp4

You can try this directly on our project page : just open a browser and play with it using your keyboard!

How we built the interactive world simulator

At a high level, Interactive World Simulator is trained in two stages. First, we trained an autoencoder that compresses RGB images into compact 2D latent representations and reconstructs them back into images. This lets the model reason in a lower-dimensional space while still producing high-fidelity pixel-level outputs.

Second, we froze this autoencoder and trained an action-conditioned dynamics model in the latent space. Given past visual latents and robot actions, the model predicts the next latent state, which is then decoded back into an image. At inference time, this process is repeated autoregressively: predicted frames become part of the context for predicting future frames. Because prediction happens in latent space and uses consistency models, the simulator can run interactively while remaining stable over long horizons.

How is it different from prior works?

This differs from several existing approaches to robot simulation and world modeling. Classical physics simulators can be powerful, but they often require manually specifying object geometry, contact dynamics, friction, and deformation, and the resulting simulation may still have a large sim-to-real gap. Recent video-generation and world-model methods offer a more data-driven alternative, but many are either not explicitly conditioned on robot actions, too slow for real-time interaction, closed-source, or unstable under long-horizon rollouts.

In contrast, our Interactive World Simulator is action-conditioned, produces physically accurate pixel-level predictions, and supports stable interactions for more than 10 minutes at 15 FPS on a single consumer RTX 4090 GPU. This makes it practical not only as a visual prediction model, but as an interactive engine for collecting policy-training data and evaluating robot policies reproducibly. It has multiple applications.

Application 1: scalable data generation

If our simulator is good enough, can demonstrations collected using this world model actually replace real demonstrations for training imitation learning policies?

To answer this, we collected demonstrations entirely inside our world simulator and then trained imitation learning policies on 0% real-world data and 100% generated data. We deployed the trained policies directly on the real robot. The deployed policies do not just complete the tasks. They remain robust under continuous human perturbations. This demonstrates that the generated data quality is comparable to real data.

https://aihub.org/wp-content/uploads/2026/07/twitter_4.mp4

Application 2: faithful policy evaluation

Evaluating a robot policy in the real world is hard: you need to reset the scene, rerun the policy many times, and compare across checkpoints under matched conditions. In practice, this is not scalable and reproducible.

In contrast, we could roll out policies in our world model for reproducible…

News is gathered automatically from public robotics & AI feeds on a schedule.

Related

Globalrobotics

Small, medium or large, a robotic fish maintains its swimming ability

The scalable robot ScaFi. 2026 CREATE Lab EPFL. CC BY SA-4.0. Propeller-powered underwater vehicles have long helped scientists explore and monitor aquatic environments. But they’re limited by their own mechanics: spinning blades can snag on vegetation, stir up sediment, and startle the wildli

SourceRobohub2 days ago
Globalrobotics

Video Friday: Life’s Better With a Little Robot Goose

Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion. IROS 2026 : 27 September–1 October 2026, PITTSBURGH CoR

SourceIEEE Spectrum4 days ago
Globalrobotics

Robot Talk Episode 163 – Robots helping people, with Aaron Edsinger

Claire chatted to Aaron Edsinger from Hello Robot about their Stretch robots that can help people with mobility issues live more independently. Aaron Edsinger is CEO and Co-Founder of Hello Robot, a robotics company dedicated to developing practical, safe, and accessible robots for people. With a pa

SourceRobohub4 days ago
Globalrobotics

An open source approach to physical AI from Intrinsic

Deborah Lupton / Pop Chips / Licenced by CC-BY 4.0 Intrinsic (an AI robotics group at Google) announced that the company is making parts of its platform open source. Intrinsic Core™ is a set of ROS-compatible capabilities for building sophisticated robotic applications. The capabilities included in

SourceRobohub5 days ago
Globalrobotics

Mexican EPICS in IEEE Team Builds Portable Educational Platform

In Guadalajara, Mexico, many high schools have motivated teachers and talented students with an interest in science, technology, engineering, and mathematics, but they lack access to advanced tools such as robotics laboratories. The resources shortfall limits the students’ opportunities for hands-on

SourceIEEE Spectrum5 days ago
Globalrobotics

What’s coming up at #IROS2026?

The 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026) will be held from 27 September – 1 October in Pittsburgh, USA. The programme includes plenary and keynote talks, workshops, tutorials, forums, competitions, and a debate. Keynote talks The keynote talks wi

SourceRobohub6 days ago