Method
Latent disturbances perturb the learned dynamics within a calibrated uncertainty set of plausible transitions, solved via efficient game-theoretic optimization.

Carnegie Mellon University
Robust decision-making in the latent space of world models, by modeling latent-space disturbances.
Latent disturbances perturb the learned dynamics within a calibrated uncertainty set of plausible transitions, solved via efficient game-theoretic optimization.
54%→15% failure rate with a robust latent safety filter and 35%→70% success rate with sample-and-verify in contact-rich manipulation.
Robust decision-making directly in the latent space of world models, validated against ground-truth robust solutions when system dynamics and disturbances are known.
Three vision-based tasks in simulation and the real world: Dubins’ car, block pouring, and egg serving with a Franka robot.
Serve a sunny-side-up egg without flipping it.

Click a trajectory to explore its imagined outcome.
Generated by the world model.
The same serving motion can place the egg on the plate or flip it, depending on unobserved friction or oil on the spatula. Robust optimization selects an action that remains effective under the worst-case disturbance in a specified set. This is a game: the disturbance maximizes the cost of the outcome, while the robot chooses an action that minimizes this worst-case cost.
A world model learns compact latent states and their dynamics from observations. It does not explicitly represent disturbances such as friction or the amount of oil. Instead, uncertainty about these factors is reflected in its distribution of predicted next states.
Key idea: model a latent-space disturbance as a perturbation to the predicted next-state distribution. Optimize this disturbance within an uncertainty set of plausible latent dynamics to induce pessimistic world-model imaginations.
Here, the disturbance picks the most adverse latent dynamics within the uncertainty set, so that the world model imagines futures in which the action is most likely to fail, while the robot chooses the action that performs best even under these pessimistic imaginations. The uncertainty set must therefore include diverse transitions that are plausible under the system, while excluding implausible ones.
Challenge: Overly pessimistic disturbances. A nearby or high-likelihood latent state need not represent a physically feasible outcome. The disturbance can exploit these model errors and imagine impossible failures, making the robot overly conservative. The uncertainty set must cover adverse transitions while constraining implausible, out-of-distribution outcomes.
Current State
Next State
Next StateKL divergence from the predicted next-state distribution
This latent-space robust optimization can be used to steer a task policy πtask at runtime, preventing hard-to-model failures under uncertainty. We instantiate it for two policy-steering paradigms, replacing nominal world-model imaginations with pessimistic yet plausible ones.
Safeguard πtask with least-restrictive filtering: evaluate the safety of the action proposed by the task policy, and intervene with the robust safety policy only when that action is doomed to fail.
The robust safety value and safety policy are learned with pessimistic imaginations induced by the latent disturbance:
Samples K candidate action sequences a(k) ∼ πtask(z), evaluates each with world-model imaginations, and executes the one with the lowest expected cost.
The learned latent disturbance generates adverse futures for each candidate, so action selection accounts for calibrated system uncertainty.
Loading trajectory…
Rollouts safeguarded by the robust safety filter. Orange marks steps where the filter overrides the task-policy action because its robust safety value becomes non-positive.
Loading trajectory…
Rollouts safeguarded by the nominal safety filter. The dotted orange line evaluates the same rollout with the robust safety value for reference.
Both videos are generated by the world model.
Nominal (left) and pessimistic (right) world-model imaginations for the same real-world observation and action.
Failure rate ↓ · n = 30
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.6333 | 0.4551–0.7813 | 19 / 30 |
| Nominal | 0.3333 | 0.1923–0.5122 | 10 / 30 |
| Ours | 0.0667 | 0.0185–0.2132 | 2 / 30 |
Failure rate ↓ · n = 30
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.9000 | 0.7438–0.9654 | 27 / 30 |
| Nominal | 0.6000 | 0.4232–0.7541 | 18 / 30 |
| Ours | 0.1333 | 0.0531–0.2968 | 4 / 30 |
Failure rate ↓ · n = 30
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.8333 | 0.6644–0.9266 | 25 / 30 |
| Nominal | 0.7000 | 0.5212–0.8334 | 21 / 30 |
| Ours | 0.2667 | 0.1418–0.4445 | 8 / 30 |
Robust rate ↑ · n = 30
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.1000 | 0.0346–0.2562 | 3 / 30 |
| Nominal | 0.1667 | 0.0734–0.3356 | 5 / 30 |
| Ours | 0.6667 | 0.4878–0.8077 | 20 / 30 |
We replay 30 failure and 20 success trajectories on three spatula surfaces while safeguarding them. The robust rate is the fraction of trajectories that remain safe across all surfaces.
Success rate ↑ · n = 20
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.1000 | 0.0279–0.3010 | 2 / 20 |
| Nominal | 0.5000 | 0.2993–0.7007 | 10 / 20 |
| Ours | 0.7500 | 0.5313–0.8881 | 15 / 20 |
Success rate ↑ · n = 20
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.4000 | 0.2188–0.6134 | 8 / 20 |
| Nominal | 0.4500 | 0.2582–0.6579 | 9 / 20 |
| Ours | 0.7000 | 0.4810–0.8545 | 14 / 20 |
Success rate ↑ · n = 20
| Method | Rate | 95% CI | Count |
|---|---|---|---|
| Base Policy | 0.4000 | 0.2188–0.6134 | 8 / 20 |
| Nominal | 0.3500 | 0.1812–0.5671 | 7 / 20 |
| UnConf. | 0.3000 | 0.1455–0.5190 | 6 / 20 |
| Ours | 0.7000 | 0.4810–0.8545 | 14 / 20 |
Left two panels: safety filtering over 20 trials with Scotch tape on the spatula. Right panel: sample-and-verify steering of π₀.₅, which evaluates 8 candidate action chunks in world-model imagination and executes the one with the best predicted outcome.
Same action · Different physics
Right video is generated by the world model.
Left: the action executed in simulation. Right: the imagined outcome of the same action, using a latent disturbance optimized without the in-distribution constraint.
Both videos are generated by the world model.
Nominal (left) and pessimistic (right) world-model imaginations for the same initial state and action.
Loading trajectory…
Both filters safeguard the same base action sequence. The plots show when each filter intervenes; the videos show whether those actions prevent failure.
2,000 rollouts per method
| Method | Rate |
|---|---|
| Base Policy | 0.4400 |
| Nominal | 0.4500 |
| Worst-of-10 | 0.4600 |
| CVaR (0.1) | 0.3900 |
| UnConf. | 0.0000 |
| Ours | 0.7300 |
Each policy is rolled out for 2,000 trajectories with randomized initial states and physics. Success: Fraction of rollouts that complete the task.
100 trajectories × 20 physics settings
| Method | Rate |
|---|---|
| Base Policy | 0.1000 |
| Nominal | 0.1100 |
| Worst-of-10 | 0.0700 |
| CVaR (0.1) | 0.1000 |
| UnConf. | 0.9000 |
| Ours | 0.7700 |
We replay 100 successful teleoperation trajectories, each under 20 different physics settings. All-safe: Fraction of trajectories safe under all 20 physics settings.
2,000 rollouts per method
| Method | Rate |
|---|---|
| Base Policy | 0.4300 |
| Nominal | 0.6000 |
| Worst-of-10 | 0.3400 |
| CVaR (0.1) | 0.0000 |
| UnConf. | 0.0000 |
| Ours | 0.8100 |
The same task with fixed physics, so uncertainty arises only from partial observability and model approximation. Success: Fraction of rollouts that complete the task.
@article{seo2026modeling,
title = {Modeling Latent Disturbances for Robust Decision-Making in World Models},
author = {Seo, Junwon and Bajcsy, Andrea},
journal = {arXiv preprint arXiv:2610.07599},
year = {2026}
}