Abstract
Temporally consistent long video generation remains a fundamental challenge. Existing methods suffer from feature drift, where entities and environments gradually change unintentionally, or content collapse, where narratives fail to progress meaningfully. We introduce A${^2}$RD, an agentic autoregressive video generation architecture that decouples creative synthesis from consistency by modeling consistency as a test-time objective. A${^2}$RD features segment-by-segment generation augmented with a novel Multimodal Video Memory (\memory{}) that tracks segment contexts and dynamics and Test-Time Scaling algorithms that verify and refine generation. For each segment, it operates in a Retrieve--Synthesize--Refine--Update (RSRU) loop: the agent retrieves relevant contexts, determines the segment generation mode (extrapolation or interpolation) adaptively, synthesizes boundary frames then video segment with refinements applied at both frame and video levels, and updates \memory{} for subsequent generation. We further develop LVbench-C, a challenging benchmark measuring long-horizon entities and environments evolving in non-linear transitions. Extensive experiments on public and LVbench-C benchmarks across one-, three-, and five-minute video generation demonstrate that A${^2}$RD generates significantly more consistent, meaningful videos than existing baselines. Human evaluations confirm strong consistency in characters, objects, and environments, with smooth motion and meaningful narrative progression.