#AAMAS2026 blue sky award winner: Foundation world models for agents in changing environments

Florent Delgrange won the Best Blue Sky Paper Award at AAMAS 2026 for his work Foundation World Models for Agents that Learn, Verify, and Adapt Reliably Beyond Static Environments . We caught up with him to find out more about his vision for agent learning.
What is the topic of your Blue Sky Ideas paper and why is it an interesting area for study?
My Blue Sky Ideas paper asks a simple but difficult question: how can an autonomous agent keep learning as its world changes without quietly losing the guarantees that made its behavior trustworthy?
Reinforcement learning and formal methods address complementary parts of this problem. Reinforcement learning allows an agent to learn by trial and error and can scale to environments for which we could never write down every rule. However, the agent is usually asked to maximize a reward. A poorly specified reward can be exploited, and a high reward does not by itself tell us that a safety or coordination requirement has been satisfied. Reactive synthesis starts from the other end: given a model of the environment and a logical description of the intended behavior, it can construct a policy that is correct by design. The difficulty is that it normally needs an explicit, fixed model, which is precisely what an agent lacks in an open and changing world.
The paper proposes a research agenda that brings these traditions into one loop. A foundation world model would be learned from experience, but structured so that a verifier can reason about it. As the agent learns a policy, it would update its model, measure the reliability of its abstraction, and check whether the policy still satisfies its specification. The verifier’s feedback could reject an unsafe update, request data from an uncertain region, or trigger a revision of the model.
This matters because real environments do not politely remain as they were during training. Goals evolve, conditions change, and, in a multi-agent system, every adapting agent changes the environment perceived by the others. Reliability therefore cannot be a certificate obtained once at deployment. It has to be maintained as the agent continues to learn.
What is your vision for foundation world models?
To me, foundation refers first to reuse, not simply to size. I do not envision a larger video predictor trained on more trajectories. I envision a persistent and structured model that an agent can carry across tasks, policies, and changing populations of agents.
One way to think about it is as a map that records more than roads. It should also tell the agent which areas have been surveyed, which conclusions depend on those areas, and when a change in the world has made an old route unreliable. A useful foundation world model should do the same for decision-making: predict what may happen, expose the structure needed for reasoning, and quantify when those predictions are trustworthy enough to support a guarantee.
I’d want such a model to have three key properties. First, it should be calibrated: every learned abstraction should come with a measure of its error or coverage, so that the agent knows where formal conclusions remain valid. Second, it should be compositional: verified local dynamics, behaviors, and certificates should be reusable when a new task is assembled. Third, it should be semantically queryable: a formal requirement or high-level instruction should help the agent derive a suitable reward model, task-specific abstraction, or policy prior with little additional experience.
The durable idea is that the world model becomes a common substrate for learning, planning, and verification. It should help an agent act, but also identify the limits of its competence, gather evidence where those limits matter, and explain why a particular behavior can or cannot currently be certified.
How does this vision differ from current models?
Most current model-based approaches focus on learning for a particular task, environment, policy, or training distribution. Their main model learning objective is predictive accuracy: reconstruct an observation, forecast the next (latent) state, or generate imagined trajectories for planning. Those are important capabilities, but average predictive metric is not the same as fitness for a particular guarantee.
Consider a model for a scenario involving an agent interacting in a warehouse, where the world model predicts almost every transition correctly but misses a rare, dangerous interaction with a forklift. Its average error may be excellent while that transition is exactly the one that determines whether a collision-avoidance claim is valid. The reverse is also possible: a compact model may ignore colors and textures yet preserve everything needed to reason about routes and collisions. For verification, the relevant question is therefore not only “How accurate is the model?” but “Which conclusions does this accuracy justify, for this policy and this requirement?”
Foundation world models would make that connection an explicit design objective. Their abstractions would carry reliability information tied to the behavior being analyzed, and this information would be revised online as the data distribution changes. Previously verified components could be reused and composed, while specifications expressed in logic or language could guide which representation and policy the agent needs for a new task.
The main difference is therefore a change in role. Current models are primarily prediction tools. The model I envisage is a persistent, analyzable basis for learning, adaptation, and formal reasoning, with language models providing a complementary semantic interface. They could translate high-level instructions into candidate specifications or propose structured model updates, while the world model grounds these proposals in experience and the verifier checks their validity. Foundation world models may benefit from scale through broader task and environment coverage, but scale…
News is gathered automatically from public robotics & AI feeds on a schedule.
Related

Small, medium or large, a robotic fish maintains its swimming ability
The scalable robot ScaFi. 2026 CREATE Lab EPFL. CC BY SA-4.0. Propeller-powered underwater vehicles have long helped scientists explore and monitor aquatic environments. But they’re limited by their own mechanics: spinning blades can snag on vegetation, stir up sediment, and startle the wildli

Video Friday: Life’s Better With a Little Robot Goose
Video Friday is your weekly selection of awesome robotics videos, collected by your friends at IEEE Spectrum robotics. We also post a weekly calendar of upcoming robotics events for the next few months. Please send us your events for inclusion. IROS 2026 : 27 September–1 October 2026, PITTSBURGH CoR

Robot Talk Episode 163 – Robots helping people, with Aaron Edsinger
Claire chatted to Aaron Edsinger from Hello Robot about their Stretch robots that can help people with mobility issues live more independently. Aaron Edsinger is CEO and Co-Founder of Hello Robot, a robotics company dedicated to developing practical, safe, and accessible robots for people. With a pa

An open source approach to physical AI from Intrinsic
Deborah Lupton / Pop Chips / Licenced by CC-BY 4.0 Intrinsic (an AI robotics group at Google) announced that the company is making parts of its platform open source. Intrinsic Core™ is a set of ROS-compatible capabilities for building sophisticated robotic applications. The capabilities included in

Mexican EPICS in IEEE Team Builds Portable Educational Platform
In Guadalajara, Mexico, many high schools have motivated teachers and talented students with an interest in science, technology, engineering, and mathematics, but they lack access to advanced tools such as robotics laboratories. The resources shortfall limits the students’ opportunities for hands-on

What’s coming up at #IROS2026?
The 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026) will be held from 27 September – 1 October in Pittsburgh, USA. The programme includes plenary and keynote talks, workshops, tutorials, forums, competitions, and a debate. Keynote talks The keynote talks wi