AI agents create virtual playgrounds to help robots get crucial training data

A new system developed at MIT uses three agents to piece together the objects, walls, and overall look of a 3D scene. Its realistic recreations of indoor spaces help robots practice skills and try out different ways of doing tasks before they’re powered on. Image: Tim Malieckal/MIT CSAIL using assets from the researchers.
By Alex Shipps
Robots walking down the street, surrounded by astounded onlookers, is an increasingly common sight. But these machines aren’t yet the do-it-all assistants you’d want working in a kitchen or factory, and a major bottleneck is data. Much like humans, robots learn best by experience. The challenge is that it’s labor-intensive and time-consuming to physically teach these machines so many actions across different settings.
“One natural idea is to use simulation as a training ground. While there has been significant progress over the last few years in the physics engines that power robotics simulators, one of the remaining challenges has been creating sufficiently rich and diverse simulation content to capture the complexity of the real world,” says Russ Tedrake, the Toyota Professor of Electrical Engineering and Computer Science (EECS), Aeronautics and Astronautics, and Mechanical Engineering at MIT, and a principal investigator at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL).
It turns out that AI agents, or semi-autonomous programs that “think” and complete well-defined tasks, could help produce the lifelike virtual settings that robots need. The new “ SceneSmith ” system developed by researchers at MIT CSAIL and Toyota Research Institute uses three agents to piece together the objects, walls, and overall look of a 3D scene. Its recreations of indoor spaces such as restaurants, bedrooms, and hotels are more realistic and detailed than prior systems, helping robots practice skills and try out different ways of doing tasks before they’re powered on. In turn, engineers save time on real-world testing.
The agents have a sense of how everyday places are supposed to look because they each call on a multi-modal system called a vision-language model (VLM), specifically the state-of-the-art VLM GPT-5.2 . It’s trained on lots of text and images from the internet to handle more visual prompts. This advanced model gives each agent a sort of spatial knowledge: First, a “designer” agent generates the elements of a scene, then a “critic” advises whether it looks realistic, and finally, an “orchestrator” manages their back-and-forth, deciding when the design is done. Once the three VLMs wrap up their creative collaboration, the scene is ready to load directly into physics simulation software.
“We’ve found that the system can construct 3D scenes the way a human designer would,” says MIT EECS PhD student Nicholas Pfaff, a CSAIL researcher and a lead author on a paper with Tedrake presenting the work. “We made over 1,300 scenes using a leading VLM that has internet-scale priors, and it made insanely creative and diverse arrangements. I hadn’t taught the system to do that in the prompts; it just improvised.”
Talk to my agent
Thanks to VLM agents, you can ask SceneSmith to do things like “generate a garage with a car, a workbench, tires stacked in the corner, and a ladder against the wall,” and get a virtual playground rich with objects a robot can tinker with. These rooms are decorated with up to six times more items per scene than prior methods, making them great for helping robots learn skills such as putting a cup in the sink, placing fruit on plates, and moving a soda can from a shelf to a table.
With so many rich virtual environments handy, you can evaluate whether your robot is ready for deployment without so much trial and error in the physical world. The researchers tested out different action plans (also called “policies”) in SceneSmith’s digital worlds, generating 100 unique spaces in the process. A VLM agent evaluated each attempt, and it found the robot’s plans were faulty, with the machine often failing at its chores. Humans agreed with the model’s verdicts over 99 percent of the time, which could help roboticists weed out flawed approaches in simulation before a robot moves in the real world.
But how realistic are these virtual worlds, really? It can be difficult to prove outright, so the researchers approached the question from several angles. The most telling test: they dropped a pretrained robot policy — an AI controller trained largely on real-world data, which had never seen a SceneSmith scene — into the generated environments. In one test, users told the system to “take the apple from the bowl and place it onto the cutting board,” and the simulated robot did exactly that. If the scenes didn’t closely resemble the real settings the policy had learned from, it simply wouldn’t have worked.
The team also teleoperated robots through the virtual spaces, guiding them to open cabinets, put away bottles, and navigate between rooms. Their experiments revealed that the environments hold up under sustained physical interaction, expanding beyond visual inspection.
Behind the scenes
The agents that SceneSmith uses each have a well-defined role in the generative process, fleshing out scenes in stages. They essentially create a floor plan and bring it to life.
Let’s say you wanted to create a scene similar to the first floor of a house. The “designer” VLM would start with a general layout, which the “critic” reviews, and then the “orchestrator” signs off. The agents repeat this approach for each step: adding furniture, placing objects on walls and then ceilings, and finally, dropping in objects that robots can manipulate. For example, the VLMs can add cabinets that the robots can open and close — an articulated item, which prior baselines didn’t often have.
At each stage, the second VLM ensures the scene is practical, advising that a bathtub is removed from a living room, for example. The third VLM ensures a high-quality scene is…
News is gathered automatically from public robotics & AI feeds on a schedule.
Related

Small, medium or large, a robotic fish maintains its swimming ability
The scalable robot ScaFi. 2026 CREATE Lab EPFL. CC BY SA-4.0. Propeller-powered underwater vehicles have long helped scientists explore and monitor aquatic environments. But they’re limited by their own mechanics: spinning blades can snag on vegetation, stir up sediment, and startle the wildli

Video Friday: Life’s Better With a Little Robot Goose
Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion. IROS 2026 : 27 September–1 October 2026, PITTSBURGH CoR

Robot Talk Episode 163 – Robots helping people, with Aaron Edsinger
Claire chatted to Aaron Edsinger from Hello Robot about their Stretch robots that can help people with mobility issues live more independently. Aaron Edsinger is CEO and Co-Founder of Hello Robot, a robotics company dedicated to developing practical, safe, and accessible robots for people. With a pa

An open source approach to physical AI from Intrinsic
Deborah Lupton / Pop Chips / Licenced by CC-BY 4.0 Intrinsic (an AI robotics group at Google) announced that the company is making parts of its platform open source. Intrinsic Core™ is a set of ROS-compatible capabilities for building sophisticated robotic applications. The capabilities included in

Mexican EPICS in IEEE Team Builds Portable Educational Platform
In Guadalajara, Mexico, many high schools have motivated teachers and talented students with an interest in science, technology, engineering, and mathematics, but they lack access to advanced tools such as robotics laboratories. The resources shortfall limits the students’ opportunities for hands-on

What’s coming up at #IROS2026?
The 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026) will be held from 27 September – 1 October in Pittsburgh, USA. The programme includes plenary and keynote talks, workshops, tutorials, forums, competitions, and a debate. Keynote talks The keynote talks wi