Skip to main content
Bibha home

Research digest,

Diffusion Controller: one control framework for diffusion model fine-tuning

Diffusion Controller, from researchers at Google and two universities, treats diffusion model fine-tuning as a control problem over the denoising process. The framework yields two reinforcement learning methods and a small side network that steers a frozen image model, and the authors report it beat LoRA in two of three training regimes on Stable Diffusion v1.4.

By Pupinder Singh, Co-Founder & CTO, Bibha AI Labs

PaperYang, T., Ryu, M., Hsu, C.-W., Tennenholtz, G., Chi, Y., Boutilier, C., & Dai, B. (2026). Diffusion Controller: Framework, algorithms and parameterization. arXiv:2603.06981. (external site)

A sealed pale stone block beside a small brass and green mechanism that bends a stream of loose grain into a smooth, ordered ribbon.
AI-generated illustration: a small controller steering a stream of noise into an ordered form beside a frozen, untouched model.

Why is steering text-to-image models so hard?

Pushing an image model harder towards a goal usually drags it away from what it learnt in pretraining, and image quality pays for that distance. Teams currently combine inference-time guidance with training-time tools such as adapters and reward-based fine-tuning. The paper argues these have been presented as separate fixes with separate objectives, so deciding how far to push has relied on trial and error rather than a shared theory.

The Google Research post illustrates the tension with a lizard wearing sunglasses. A lightly steered model may draw a convincing lizard and forget the sunglasses. A heavily steered one may include them but warp the animal's face. Either way, the user receives an image that misses part of the brief.

Access adds a second constraint. Popular adapters such as LoRA insert trainable layers inside the network, which requires white-box access to the weights. The paper describes a common middle case, grey-box access, where a proprietary or safety-sensitive model stays sealed but exposes limited signals such as its per-step noise predictions or denoising trajectory. Methods that need the weights cannot be used there.

How does Diffusion Controller reframe diffusion model fine-tuning?

It treats the denoising process as a stochastic control problem. In the authors' formulation, a controller reweights the pretrained model's transition from one noisy state to the next, a reward scores only the finished image, and a divergence penalty charges for every step that strays from the original behaviour. Fine-tuning then means finding the control that best balances reward against that penalty.

The mathematical setting is a linearly solvable Markov decision process, a class of control problems in which the controller acts directly on transition probabilities rather than choosing separate actions. Earlier reinforcement learning methods for diffusion, such as DDPO and DPOK, modelled the next noisy image as an action. The authors point out that this action is really just the next state, so shaping the transition itself is the more direct description.

The penalty is a general f-divergence. Choosing the Kullback-Leibler divergence recovers the classic version of this control problem, and a regularisation coefficient sets how strongly the model is held near its pretrained behaviour. Because the ideal controlled transitions are expensive to sample from directly, the authors instead train an ordinary diffusion denoiser whose final outputs follow the same target distribution.

Two fine-tuning methods from one framework

The control view produces two reinforcement learning recipes that need nothing more than a reward model applied to the final image. Both come from the same optimality conditions, which is the heart of the paper's unification claim: methods developed separately turn out to be variations on one objective.

The same model structure can also be trained with plain supervised fine-tuning when example images from the target distribution exist, so one design covers both data-driven and reward-driven adaptation.

  • Policy gradient with a PPO-style update. The model improves through small, repeated updates weighted by an advantage estimate. A clipping rule caps how far any single update can move the model, which keeps training stable. With no regularisation the update reduces to DDPO, and with a Kullback-Leibler penalty it closely resembles DPOK.
  • Reward-weighted loss. The standard denoising loss is reweighted by each sample's reward, which turns an intractable objective into an ordinary regression. Under the Kullback-Leibler penalty the authors prove this loss shares its minimiser with the ideal objective. Exponential weighting follows from that case, and polynomial weighting from an alpha-divergence penalty.

What does the side network add to a frozen model?

It supplies the correction while the original model stays untouched. The control analysis shows that the optimal fine-tuned predictor equals the pretrained predictor plus a small, structured adjustment. Diffusion Controller therefore freezes the backbone and trains a separate lightweight network that reads the backbone's intermediate output at each step, together with the prompt and timestep, and returns that adjustment.

The input matters. Rather than the noisy latent itself, the side network receives the reverse mean, the backbone's estimate of where the next denoising step should land. Its output splits into a single gating value and a correction term, mirroring the decomposition in the paper's main proposition, which the authors relate to cross-attention over features of the reverse mean. The implementation is a small UNet that runs in latent space.

Two practical details stand out. The final layer starts at zero, so before training the combined model behaves exactly like the pretrained one, much as a fresh LoRA adapter does. A strength setting on the side network can also be changed at inference time, letting a user turn the steering up or down without retraining.

How was Diffusion Controller evaluated?

All experiments used Stable Diffusion v1.4, with the text encoder and image autoencoder frozen and only the latent denoiser adapted. The authors ran three regimes: supervised fine-tuning on preferred images from Human Preference Dataset v2, plus reward-weighted and PPO fine-tuning against the Human Preference Score v2 reward model. The headline metric is how often a model beats the pretrained one under that score.

Supervised runs used 1000 iterations at a batch size of 128. Reward-weighted runs used 2000 iterations and PPO runs 2400 iterations, both at a batch size of 64, and classifier-free guidance was fixed at 7.5. The authors also tracked CLIP, CLIP-Aesthetics and PickScore to check that other qualities held, and recruited more than 50 paid raters to compare PPO models on 50 prompts.

  • DiffCon, grey-box. The proposed side network, fed the reverse mean and combined through the gated correction. It trains about 12 million parameters.
  • DiffCon-Naive, grey-box baseline. A near-identical side network that sees the noisy latent and adds an ungated correction, isolating the effect of the input choice and output structure.
  • LoRA, white-box baseline. Rank 16 adapters inserted into the backbone, with about 17 million trainable parameters.
  • DiffCon-J, white-box. Rank 4 LoRA adapters and the side network trained together, about 16 million parameters in total.
  • DiffCon-S, white-box. The same two components trained separately and combined only at evaluation.

What results do the authors report?

In two of the three regimes the grey-box controller beat LoRA while training fewer parameters. The authors report win rates against the pretrained model of 0.6667 for DiffCon versus 0.5766 for LoRA under supervised fine-tuning, and 0.6815 versus 0.6109 under reward-weighted fine-tuning. Under PPO the picture changes: LoRA reached 0.9048 and the grey-box DiffCon 0.6957.

The white-box combinations performed best under PPO, at 0.9353 for DiffCon-J and 0.9315 for DiffCon-S. The Google Research post summarises this fully unlocked setting as a 90% win rate over the base model. DiffCon-S also recorded the highest reward-weighted result, 0.7091.

The naive side network barely moved, with win rates of 0.5655, 0.5060 and 0.5201 across the three regimes. Because it shares the controller's size and backbone, the gap suggests the gains come from what the side network sees and how its output is combined, not from extra capacity alone.

Secondary checks were steady. CLIP, CLIP-Aesthetics and PickScore stayed close to the pretrained model's values, which the authors read as no loss of other qualities. In the human study on PPO models, they report each of their variants beating its own baseline. The paper's polynomial reward weighting also outperformed the reward-weighted loss used in DPOK while drawing a quarter of the samples.

Limitations and open questions

The evidence is promising but narrow. Every result comes from one backbone, Stable Diffusion v1.4, which predates systems the paper itself cites such as Stable Diffusion 3 and Flux.1, and from one preference signal, Human Preference Score v2. The authors name personalisation, safety alignment and transfer learning as future directions; none is tested in the paper.

Our own reading of the paper adds four cautions.

  • Reward and metric overlap. The reward-driven runs optimise Human Preference Score v2 and are judged mainly by the same score. The secondary metrics and human study reduce, but do not remove, the risk that models learn to please the scorer.
  • Best-setting reporting. For each checkpoint the authors test several side-network strength settings and report the best win rate. A single fixed setting would likely score lower.
  • Grey-box is not black-box. The controller needs the backbone's intermediate outputs at every denoising step. Many hosted image services return only finished images, so the method fits vendors that expose those signals, not every closed model.
  • Uneven coverage. The human evaluation covers only PPO models and compares each variant with its own baseline, not the grey-box controller with LoRA. Neither the paper nor the post mentions a code release.

Why does this matter for enterprise image generation?

In our analysis, the idea worth noting is separation. A vetted base model kept frozen, plus a small controller trained for a specific goal, is easier to version, test and roll back than a fully modified copy of the network. A runtime strength setting adds a simple way to trade prompt adherence against fidelity for each use case.

The practical catch is access. The approach needs a model provider willing to expose intermediate denoising outputs, and the reward decides what improves. A team adapting images to brand style or product accuracy would need a reward that measures exactly that, because results on Stable Diffusion v1.4 with a general preference score do not transfer automatically to such goals.

The paper also offers a useful vocabulary. Framing guidance, adapters and reward fine-tuning as variations on one control problem gives engineering teams a clearer way to compare methods and to reason about the cost of steering further from a base model.

Is Diffusion Controller available to use?

Not as released software. Neither the paper nor the Google Research post links to code or trained weights. The paper, arXiv:2603.06981, is distributed under arXiv's standard non-exclusive licence, which is why this digest reproduces none of its figures. The preference dataset and reward model behind the experiments, Human Preference Score v2, are public from their own authors.

Questions and answers

What is Diffusion Controller?

Diffusion Controller, or DiffCon, is a framework from researchers at Google Research, Google DeepMind, Carnegie Mellon University and Yale University. It describes diffusion model fine-tuning as optimal control of the denoising process, derives policy-gradient and reward-weighted training methods from that view, and proposes a small side network that steers a frozen backbone. The paper tests it on Stable Diffusion v1.4.

How is Diffusion Controller different from LoRA?

LoRA inserts small trainable matrices inside the model, so it needs access to the weights. Diffusion Controller leaves the model untouched and trains a separate network that reads the model's intermediate predictions and adds a correction. In the authors' tests the side network beat LoRA in supervised and reward-weighted fine-tuning with fewer parameters, while LoRA stayed ahead under PPO. The two can also be combined.

Can Diffusion Controller fine-tune a closed image model?

Only a partly closed one. The method keeps the backbone frozen, but it needs the model's intermediate reverse mean at each denoising step, which the paper calls grey-box access. A provider that returns only finished images would not supply that signal, so the method suits settings where a vendor exposes per-step outputs or runs the controller alongside its own model.

What reward and benchmark did the experiments use?

The reward-driven runs optimised Human Preference Score v2, a model trained to predict which images people prefer for a prompt, and the supervised runs trained on preferred images from the matching dataset. Results are reported as win rates against the unmodified Stable Diffusion v1.4 on the benchmark's test prompts, with CLIP, CLIP-Aesthetics, PickScore and a human study as secondary checks.

References

  1. Yang, T., Ryu, M., Hsu, C.-W., Tennenholtz, G., Chi, Y., Boutilier, C., & Dai, B. (2026). Diffusion Controller: Framework, algorithms and parameterization. arXiv:2603.06981. https://arxiv.org/abs/2603.06981 (external site)
  2. Wu, X., Hao, Y., Sun, K., Chen, Y., Zhu, F., Zhao, R., & Li, H. (2023). Human Preference Score v2: A solid benchmark for evaluating human preferences of text-to-image synthesis. arXiv:2306.09341. https://arxiv.org/abs/2306.09341 (external site)
  3. Rombach, R., Blattmann, A., Lorenz, D., Esser, P., & Ommer, B. (2021). High-resolution image synthesis with latent diffusion models. arXiv:2112.10752. https://arxiv.org/abs/2112.10752 (external site)
  4. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-rank adaptation of large language models. arXiv:2106.09685. https://arxiv.org/abs/2106.09685 (external site)
  5. Ho, J., & Salimans, T. (2022). Classifier-free diffusion guidance. arXiv:2207.12598. https://arxiv.org/abs/2207.12598 (external site)
  6. tgxs002. (n.d.). HPSv2: Human Preference Score v2 benchmark and reward model [Computer software]. GitHub. https://github.com/tgxs002/HPSv2 (external site)

Original article

Hsu, C.-W., & Ryu, M. (2026, 29 September). How Diffusion Controller unifies and simplifies AI image generation. Google Research Blog. https://research.google/blog/how-diffusion-controller-unifies-and-simplifies-ai-image-generation/ (external site)

This is Bibha's independent summary of published research. Bibha is not affiliated with the authors or Google.

All news and research