Abstract
Bidirectional video diffusion models denoise entire videos in parallel, yet when trained on effectively unlimited in-distribution data from procedural generators, continue to violate physical laws and simple symbolic rules. We introduce Serial-to-Parallel Diffusion (S2PD), which performs autoregressive diffusion at high noise before switching to parallel diffusion at low noise. The autoregressive phase provides the serial computation needed to coordinate interdependent events and produce valid state transitions while the parallel phase jointly refines the entire video and reduces sampling time relative to fully serial generation. We implement S2PD with two architectures: a pixel-space diffusion transformer trained from scratch and a pretrained video model adapted through LoRA fine-tuning with causal attention. Across games, physical simulations, and real video, S2PD follows rules more reliably than matched bidirectional baselines and generates videos with greater temporal stability and sampling efficiency than other serial methods.
Overview
Results
All serial methods greatly outperform the bidirectional diffusion baseline across all datasets. They also generate multiple plausible continuations given a single start image. Our method retains the benefits of serial generation while achieving better temporal stability and sampling efficiency than other serial methods.
Serial outperforms parallel generation.
One image, multiple continuations.
Improved temporal stability.
Video Gallery
We train a dedicated model for each dataset and method and visualize uncurated validation samples below. All serial methods perform similarly to each other and greatly outperform the bidirectional baseline. Red overlays visualize localized errors on the video. Use the toggle above to turn these error diagnostics on/off.
Discrete datasets
Continuous datasets
Video datasets
Citation
@misc{hu2027s2pd,
title = {S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation},
author = {Hu, Jeffrey and Olmeda Reino, Daniel and Tewari, Ayush},
year = {2027},
note = {Manuscript in preparation}
}
Acknowledgements
This work was funded in part by Toyota Motor Europe. We acknowledge resources provided by the Isambard-AI National AI Research Resource.