Skip to main content
Bibha home

Research digest,

How agents keep long-form video generation coherent: Co-Director, CANVAS, A²RD and VQQA

Researchers at Google and partner universities have published four agentic frameworks that target drift and cascading errors in long-form video generation. Co-Director plans the creative direction, CANVAS tracks the story world, A²RD builds minutes-long video segment by segment, and VQQA repairs clips by questioning them.

By Om Parkash, Principal & Head of Engineering, Bibha AI Labs

PaperSong, Y., Song, Y., Losier, N., Hodson, N., Jin, Y., Zhu, R., Xu, Y., Vlasic, D., Claassen, C., Leon, J., LeViet, K. G., Chomyn, Z., Timmons, J., Slatkin, B., Penberthy, S., & Pfister, T. (2026). Co-Director: Agentic generative video storytelling. arXiv:2604.24842. (external site)

A curved row of clear glass panels, each holding the same terracotta vase under a stone arch, linked by green threads to one dark green sphere.
AI-generated illustration: a sequence of frames that keep the same scene consistent, all tied to a shared memory.

Why is long-form video generation still hard?

Video diffusion models can render a convincing clip lasting seconds, but a story needs many clips that agree with each other. Across shots, a character's clothing shifts, rooms rearrange and props appear or vanish. Agentic pipelines that chain a script writer, an image model and a video model help, yet each stage is prompted on its own, so an early mistake carries into everything downstream.

The authors describe this as a credit assignment problem: when a finished video fails, it is hard to tell which upstream prompt caused it. They also name two failures that grow with length. In feature drift, people and settings change for no narrative reason; in content collapse, the story stops moving forward. Repairing both by hand is what makes long AI video slow and costly.

The post answers with four frameworks, each aimed at one bottleneck and each acting as an orchestration layer above existing models rather than as a new video generator.

Four frameworks, four bottlenecks

Each framework handles a different stage of production, from deciding what a video should say to repairing a single faulty clip. The post presents them as complementary pillars rather than one fixed pipeline, and each has its own paper, benchmarks and results.

According to the post, running as a layer over Gemini and Veo means outputs keep those models' safety measures, including SynthID watermarking, and extra classifiers can check the assembled video, since individually safe clips can combine in unintended ways.

  • Co-Director. Plans the creative direction for a whole video and searches for the configuration that scores best with a multimodal judge. It is due to appear at COLM 2026.
  • CANVAS. Keeps an explicit record of characters, locations and object states so storyboard frames stay consistent across shots. It is due to appear at EMNLP 2026.
  • A²RD. Generates long video one segment at a time, using a multimodal memory and self-checks to hold identity and setting steady over minutes.
  • VQQA. Asks targeted visual questions about a generated clip and uses the answers to rewrite the prompt, fixing flaws without touching model weights.

How does the AI video co-director plan a whole video?

It treats a video as one optimisation problem instead of a chain of prompts. An orchestrator agent picks a creative configuration on three axes, creative strategy, narrative mode and aesthetic archetype, using a multi-armed bandit. That choice is written into the system prompt of every downstream agent, so script, keyframes, clips and audio all serve the same intent.

Pre-production agents expand the brief into a storyline, any missing visual assets and a storyboard. A keyframe agent then fixes how characters and scenes look before a video agent adds motion and an audio agent adds voiceover and music. Storylines and keyframes are scored and regenerated locally when they fall short, and keyframes are judged as a set so that mismatches between them surface.

A multimodal judge then scores the finished cut separately on each of the three axes. Feeding these factored scores back to the bandit, which runs the UCB1 algorithm, tells the system which choice helped or hurt, sidestepping the credit assignment problem. A language model warm start steers early choices towards sensible settings, because blind exploration is expensive when every trial means rendering a video.

The paper applies the framework to advertising, which demands storytelling and hard constraints at once: keeping product and logo assets faithful, matching an audience and conveying a value proposition. The evaluated videos are short, four shots totalling 12 seconds, so the minutes-long results in this suite come from A²RD rather than Co-Director.

How does CANVAS keep characters and places consistent?

By tracking the story world explicitly. Before drawing anything, a planner reads the shot list and records which outfit each character wears in each shot, which shots share a location and how objects change state, such as an artefact moving from present to stolen to absent. Image generation then follows that plan shot by shot.

A persistent visual memory stores character anchors, location backgrounds and recent frames. For a direct continuation, CANVAS conditions on the previous frame. For a return to a known place from a new angle, it retrieves the stored background. For a new place, it starts fresh and saves what it creates. Several candidate frames are drawn per shot and scored with questions derived from the plan, and memory is updated after each choice.

CANVAS outputs storyboards, a sequence of keyframes, using a Gemini image model as its generator. It needs no training, and the authors compare it with other training-free methods that they re-implemented on the same backbone.

Diagram of the CANVAS storyboard pipeline with a global continuity plan, a visual memory of characters and locations, and candidate frame selection.
CANVAS works in two stages: a global plan of outfits, locations and object states, then shot-by-shot retrieval from visual memory, candidate generation, constraint-aware selection and a memory update.Figure 2 from Mondal et al. (2026), arXiv:2604.13452, licensed CC BY 4.0. Converted to WebP.

How does A²RD generate minutes-long video?

It builds the video one segment at a time in a closed loop. For each segment, the agent pulls relevant context from a multimodal memory, decides how to generate, creates boundary frames and then the clip, checks and refines both, and writes the result back to memory. The authors call this a retrieve, synthesise, refine and update cycle, and it requires no training.

The memory keeps three records per segment: text describing entities, their changes, spatial relations and camera movement; start and end keyframes; and the clip itself. Before the first segment, the agent also creates global reference images for each character and setting, ordering them so that, for example, a person is drawn after the room they depend on.

The central decision is adaptive. When the next segment stays in the same place, A²RD extrapolates from the start frame so the story can move naturally. When the story jumps to a new or earlier setting, it interpolates between planned start and end frames to pin appearance down. Frames and clips are scored against rubrics of eight and ten criteria, then edited or regenerated, with a prompt optimiser that learns from earlier successful and failed fixes.

A parallel variant, A²RD-Par, creates all boundary frames first and then renders the clips at the same time. The authors report it keeps characters consistent but loses some environmental consistency and smoothness, a trade made for speed.

Architecture diagram of A²RD showing inputs, memory initialisation, a multimodal video memory and an agentic segment generation pipeline.
In A²RD, inputs seed a multimodal video memory, and each segment passes through retrieval, an adaptive choice of extrapolation or interpolation, and self-improvement loops for frames and video before memory is updated.Figure 4 from Long et al. (2026), arXiv:2605.06924, licensed CC BY 4.0. Converted to WebP.

How does VQQA repair a video without access to model weights?

It edits the prompt, not the pixels. A question generation agent writes visual questions tailored to the prompt and any reference images, covering prompt alignment, visual quality and fidelity to the conditions. A question answering agent scores the video on each question from 0 to 100, and a refinement agent rewrites the prompt to address the weakest answers before the generator tries again.

The authors call these critiques semantic gradients: natural language feedback that does the job numerical gradients do in training, but through a black-box interface. Because repeated fixes can drift from what the user first asked for, a separate global rater scores every candidate against the original prompt and keeps the best one. The loop stops once a candidate reaches the target score or progress stalls.

The post gives two examples: a box-shaped balloon that the base model renders as a rigid, textureless cube, and a duet in which one performer suddenly plays the other's instrument after a cut. In both cases, the revised prompt led the generator to a consistent result.

Diagram of the VQQA framework with question generation, question answering and prompt refinement agents and a global rater choosing the final video.
The VQQA loop: agents generate and answer visual questions to build a score report, a refinement agent rewrites the prompt, and a global rater picks the best video from all candidates against the original request.Figure 3 from Song et al. (2026), arXiv:2603.12310, licensed CC BY 4.0. Converted to WebP.

What results do the authors report?

Each paper reports gains over its chosen baselines, mostly on automatic metrics scored by multimodal models, with human studies as a cross-check. The figures below are the authors' own. The benchmarks differ, so the numbers cannot be compared across frameworks.

Three of the benchmarks are new, built by the teams to probe failures that existing tests rarely reach, and all three use fictional or model-generated material. GenAD-Bench pairs 200 fictional products from 50 fictional brands with a stereotypical and an unconventional audience, giving 400 advertising scenarios; the authors chose fictional brands to avoid copyright conflicts and memorised training data. HardContinuityBench, produced with GPT-5.2, packs storyboards with long gaps before scenes return, costume changes and props that change state, and the authors acknowledge it is small. LVbench-C holds text-only scenarios for three-, five- and ten-minute videos in which characters, objects and settings vanish for at least 10 segments before returning.

  • Co-Director. On GenAD-Bench it averaged 81.4 out of 100, against 75.7 for random search over the same creative options, 68.5 for its own unoptimised pipeline and 63.6 for Veo 3.1 alone. In a human study, mean opinion scores on a scale of 1 to 5 were 3.96 for Co-Director and 3.71 for Veo 3.1.
  • CANVAS. Against the best-performing baseline, the authors report consistency gains of 21.6% for backgrounds, 9.6% for characters and 7.6% for props. Annotators preferred CANVAS in up to 90.9% of pairwise comparisons against AutoStudio and 68.3% against the strongest baseline.
  • A²RD. On one-minute VBench-Long videos it reached narrative coherence of 0.90 against 0.75 for the best baseline, and character consistency of 0.74 against 0.57. Human raters gave it 4.68 out of 5 on average, against 3.93 for VideoMemory. On ten-minute scenarios, a model judge put its character consistency at 90.5%, objects at 91.5% and environments, the hardest case, at 84.0%.
  • VQQA. With the open CogVideoX-5B generator, it raised the T2V-CompBench average from 41.89% to 53.46%, an absolute gain of 11.57 points, and improved VBench2 by 8.43 points. With Veo 3.1 it lifted the same average from 55.93% to 61.81%. In the CogVideoX-5B runs on T2V-CompBench it stopped after 1.245 refinement rounds on average, about 7.23 vision-language model calls per prompt.

Limitations and open questions

The papers name several limits themselves. A²RD may generate six videos and six images per segment and uses more compute than its baselines. VQQA's sequential loop is slower than parallel best-of-N sampling. CANVAS becomes expensive when it draws several candidates per shot, and it does not yet model fine-grained physical interaction or complex motion.

Our own reading adds four cautions.

  • Model judges grade most results. Co-Director, CANVAS and A²RD lean heavily on multimodal models as judges. The authors validate them against people, and Co-Director's judge comes close to human agreement on narrative criteria but follows visual quality less closely.
  • New, small or self-built benchmarks. Several headline numbers come from benchmarks the authors created. That is reasonable for a new problem, but it makes independent replication important.
  • Short versus long. Co-Director's evaluated videos run 12 seconds. The minutes-long evidence rests on A²RD, whose ten-minute results come from a model judge on ten scenarios.
  • Bounded by the base models. VQQA cannot fix what the generator cannot draw, and A²RD needs component models that follow instructions well. The CANVAS authors also warn that better consistency could make misleading synthetic video easier to produce.

Why does this matter for enterprise video work?

In our analysis, the pattern matters more than any single score. All four frameworks wrap planning, memory and critique around generators the user does not control, which is the position most enterprises are in with hosted video models. Explicit state, such as a record of who wears what in which scene, also makes failures easier to inspect than one opaque prompt.

The costs are real. More generator calls per segment raise spend and latency, and model judges need calibrating against the reviewers who will approve the work. For marketing, training or product video, a realistic near-term aim is fewer manual fixes on short sequences, with longer narratives still under human direction, which the authors say they intend to keep in creators' hands.

Is the code available?

Partly. Co-Director's implementation sits in the public genmedia-izumi-agent repository from Google Cloud Platform, and its project page lists data as coming soon. A²RD has a public repository under the MIT licence that currently holds project materials, with the LVbench-C dataset marked as coming soon. The CANVAS project page says code is coming soon, and VQQA has a project page with examples.

All four papers are on arXiv under CC BY 4.0, which allows the credited figure reuse in this digest.

Questions and answers

What is the AI video co-director?

It is the name the Google Research post gives to a suite of agentic frameworks for long video, and to the Co-Director paper at its centre. Co-Director uses a multi-armed bandit to choose a creative strategy, narrative mode and aesthetic style, passes that choice to every production agent, and learns from a multimodal judge's scores which choices worked.

How do these systems keep a character consistent across shots?

They store and reuse explicit references. CANVAS keeps anchor images for each character's outfit at each point in the story and retrieves them for every shot. A²RD keeps text descriptions, keyframes and clips of earlier segments and creates global reference images before it starts. Both check new frames against those references and regenerate when a check fails.

Does VQQA need access to the video model's weights?

No. VQQA works only through prompts. It asks a vision-language model targeted questions about a generated clip, turns low scores into prompt revisions and asks the generator again, then picks the best candidate against the original request. The authors report that it worked with both the open CogVideoX-5B model and Veo 3.1.

How long are the videos these frameworks produce?

It depends on the framework. Co-Director's evaluated advertisements are four shots and 12 seconds long. A²RD was tested on videos of about one, three and five minutes, plus ten-minute scenarios, and the Google Research post shows a ten-minute film made with it. CANVAS produces storyboard keyframes rather than finished video.

Can businesses use these frameworks today?

Not as finished products. Co-Director's research code is public, A²RD and CANVAS list data or code as coming soon, and all four depend on capable hosted models such as Gemini and Veo. Teams can study the published designs, but should expect substantial engineering, compute cost and their own evaluation before relying on them for production video.

References

  1. Song, Y., Song, Y., Losier, N., Hodson, N., Jin, Y., Zhu, R., Xu, Y., Vlasic, D., Claassen, C., Leon, J., LeViet, K. G., Chomyn, Z., Timmons, J., Slatkin, B., Penberthy, S., & Pfister, T. (2026). Co-Director: Agentic generative video storytelling. arXiv:2604.24842. https://arxiv.org/abs/2604.24842 (external site)
  2. Mondal, I., Song, Y., Parmar, M., Goyal, P., Boyd-Graber, J., Pfister, T., & Song, Y. (2026). CANVAS: Continuity-aware narratives via visual agentic storyboarding. arXiv:2604.13452. https://arxiv.org/abs/2604.13452 (external site)
  3. Long, D. X., Song, Y., Kan, M.-Y., Pfister, T., & Le, L. T. (2026). A²RD: Agentic autoregressive diffusion for long video consistency. arXiv:2605.06924. https://arxiv.org/abs/2605.06924 (external site)
  4. Song, Y., Pfister, T., & Song, Y. (2026). VQQA: An agentic approach for video evaluation and quality improvement. arXiv:2603.12310. https://arxiv.org/abs/2603.12310 (external site)
  5. Google Cloud Platform. (n.d.). Ads Co-Director, genmedia-izumi-agent [Computer software]. GitHub. https://github.com/GoogleCloudPlatform/genmedia-izumi-agent/tree/main/demos/backend/ads_codirector (external site)
  6. Long, D. X. (n.d.). AARD: A²RD project repository [Computer software]. GitHub. https://github.com/dxlong2000/AARD (external site)

Original article

Song, Y., & Song, Y. (2026, 24 September). Automating coherent long-form video generation. Google Research Blog. https://research.google/blog/coherent-long-form-video-generation/ (external site)

This is Bibha's independent summary of published research. Bibha is not affiliated with the authors or Google.

All news and research