A world model should reproduce what the actions determine, and invent only what they don't. D3WM learns where that boundary is.
Abstract
World models of visual information predict future image frames conditioned on given action sequences. This predicted scene evolution is driven by deterministic changes such as controlled motion on a robot platform, as well as stochastic changes such as rendering of previously occluded regions. Our key insight is that the spatial location of these stochastic changes is itself learnable.
We leverage this insight by jointly regressing both the deterministic next-state prediction and the spatial uncertainty map, then using the uncertainty map to guide a conditional diffusion process. We propose the Decoupled Drift-Diffusion World Model (D3WM), which uses a transformer trained with a Gaussian negative-log-likelihood objective to produce both the deterministic next-state prediction and a spatial map of heteroscedastic aleatoric uncertainty. The learned uncertainty map then provides a structured spatial gate for a conditional diffusion model: deterministic regions are anchored to the predicted state, while uncertain regions, such as disocclusions or dynamic textures, are sampled by the diffusion process. We frame this decomposition under a variance-gated stochastic differential equation that separates deterministic drift from stochastic diffusion. D3WM shows improved results over competitive baselines in two distinct types of action spaces: trajectory planning for robotic manipulation from generated images, and camera movement in long-horizon novel view synthesis.
Method
A variance-gated SDE over the observation axis
We write the transition from a start latent \(z_0\) under an action sequence \(a_{[0,t]}\) as a stochastic differential equation over the observation time axis, which names the two parts separately:
Drift — regress the integrated effect, not the rate
Rather than learning the instantaneous coefficient \(f_\theta\) and integrating it numerically, we regress its integrated effect over \([0,t]\) in a single forward pass. A deterministic transformer returns the conditional mean together with a spatial map of heteroscedastic aleatoric variance:
Training by Gaussian negative log-likelihood forces \(\sigma_t^{2}\) to be large exactly where the deterministic prediction is ambiguous — disocclusions, dynamic texture:
Diffusion — a learned sampler in place of the Wiener term
Integrating Eq. 1 by Euler–Maruyama would draw the stochastic term from an isotropic Gaussian, which fills disoccluded regions with static rather than plausible texture. We replace that term with a conditional ControlNet \(g_\theta\), conditioned on \(\mu_t\), \(\sigma_t^{2}\) and the reference frame \(z_0\) — an image-manifold-aware sampler for the same next-state distribution:
Variance-gated inference
At sampling time \(\sigma_t^{2}\) becomes a spatial gate rather than only a conditioning signal. With a diffusion-time anchoring schedule \(\lambda_\tau \in [0,1]\), define a per-pixel trust map, and at each reverse step \(\tau\) replace the diffusion model's latent with a weighted combination of it and a forward-noised version of the deterministic prediction:
Both terms are valid noisy latents at noise level \(\tau\), so the substitution stays compatible with the next reverse step. Where \(\sigma_t^{2}\!\to\!0\) the conditional latent collapses toward \(\mu_t\) and geometric structure is preserved; where \(\sigma_t^{2}\!\to\!1\) the pixel is handed to the diffusion model and new content is synthesized. The model is deterministic where the actions decide the answer, and generative only where they do not.
Results — Real PushT
Predicted frames good enough to plan on
We predict 16 frames ahead from a fixed and a wrist camera, then hand the predicted frame to a pre-trained PushT Diffusion Policy and let it plan. The question is not how the image looks — it is whether a policy can act on it.
- 3.5× lower trajectory error than iVideoGPT: 12.86 ADE against 45.13.
- Gripper and object land where the actions put them, instead of the hallucinated grippers and deformed T that iVideoGPT produces.
Trajectory error on Real PushTTable 1
| Method | ADE ↓ | FDE ↓ |
|---|---|---|
| Ground truth frame | 1.552 | 1.319 |
| iVideoGPT | 45.134 | 65.228 |
| D3WM | 12.864 | 17.550 |
<video autoplay loop muted playsinline> element.Results — BridgeData V2
Predict, rank, and act
Tabletop manipulation from a fixed camera, against IRASim and WorldGym — measured on the three jobs a world model actually has, not just on how well it reconstructs a frame.
- Best open-loop prediction at every horizon: μt leads PSNR by +1.4 to +1.6 dB over IRASim, with the full model second.
- Best action verifier. Ranking candidate action chunks, D3WM closes 38–40% of the random-to-oracle gap against IRASim's 23–26%, at half the compute per candidate.
- 4.9 ms per candidate for the deterministic pass — 204 candidates a second against IRASim's 0.2, the only variant cheap enough to screen at control rate.
- Driving a frozen Octo policy, the stochastic component is worth 1.063 ADE against 1.445 for the deterministic prediction alone. IRASim still leads here, since it diffuses a whole trajectory where we predict one frame.
Open-loop predictionTable 2
| Method | PSNR ↑ | LPIPS ↓ | ||||||
|---|---|---|---|---|---|---|---|---|
| 1 | 2 | 3 | 4 | 1 | 2 | 3 | 4 | |
| copy-frame (control) | 24.48 | 21.48 | 20.04 | 19.19 | .051 | .084 | .105 | .121 |
| IRASim | 23.34 | 22.37 | 21.64 | 21.12 | .101 | .116 | .127 | .136 |
| WorldGym | 23.80 | 21.10 | 19.70 | 18.86 | .059 | .093 | .118 | .137 |
| D3WM μt | 24.73 | 23.54 | 23.09 | 22.76 | .086 | .099 | .107 | .112 |
| D3WM | 24.29 | 22.80 | 22.32 | 21.91 | .104 | .118 | .124 | .130 |
Candidate verificationTable 3
| Verifier | ms/cand | K@1s | MSE @ 4 | 16 | 32 |
|---|---|---|---|---|---|
| random | — | — | .040 | .038 | .038 |
| oracle | — | — | .006 | .003 | .003 |
| IRASim | 5478 | 0.2 | .031 | .029 | .030 |
| WorldGym | 1094 | 0.9 | .044 | .056 | .063 |
| D3WM μt | 4.9 | 204 | .054 | .043 | .038 |
| D3WM | 2643 | 0.4 | .027 | .024 | .022 |
Policy consistencyTable 4
| World model | ADE ↓ | FDE ↓ |
|---|---|---|
| copy-frame (control) | 1.577 | 1.994 |
| WorldGym | 1.448 | 1.783 |
| D3WM μt | 1.445 | 1.649 |
| D3WM | 1.063 | 1.193 |
| IRASim | 0.846 | 1.035 |
Results — Novel view synthesis
Extrapolating 12 frames from two views
Camera movement as the action: two input views on the DL3DV benchmark, extrapolating 3 to 12 frames, then the same model transferred to RealEstate10K with no retraining.
- Beats DepthSplat on PSNR at every horizon, and on all three metrics by 9 frames out — the gap widens with distance as splatting's black holes expand and ours are filled.
- +3.1 dB PSNR and 39% lower LPIPS than NWM on RealEstate10K, a dataset we never trained on.
DL3DV, in-distributionTable 5
| Extrap. | Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|
| 3 fr. | DepthSplat | 20.87 | 0.761 | 0.200 |
| D3WM | 21.48 | 0.669 | 0.222 | |
| 6 fr. | DepthSplat | 17.39 | 0.643 | 0.298 |
| D3WM | 20.12 | 0.609 | 0.283 | |
| 9 fr. | DepthSplat | 15.40 | 0.557 | 0.371 |
| D3WM | 19.19 | 0.563 | 0.338 | |
| 12 fr. | DepthSplat | 14.14 | 0.496 | 0.421 |
| D3WM | 18.35 | 0.525 | 0.382 |
RealEstate10K, zero-shotTable 6
| Extrap. | Method | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
|---|---|---|---|---|
| 3 fr. | NWM | 18.45 | 0.632 | 0.233 |
| D3WM | 21.55 | 0.790 | 0.142 | |
| 6 fr. | NWM | 15.86 | 0.554 | 0.315 |
| D3WM | 19.21 | 0.619 | 0.189 | |
| 9 fr. | NWM | 14.45 | 0.509 | 0.375 |
| D3WM | 17.63 | 0.556 | 0.237 | |
| 12 fr. | NWM | 13.57 | 0.480 | 0.422 |
| D3WM | 16.78 | 0.522 | 0.266 |
Citation
BibTeX
@article{
}