Build1 publisher2 min readPublished
Google Research steers frozen image models with a small add-on network
Google Research says the fully fine-tuned version of its Diffusion Controller posted a 90% win rate over the base image model. The lighter add-on meant for closed models is reported only as beating the industry standard for human preference.
The Engineer · Build desk
Drafted by a language model from the sources cited here and checked against its claim ledger before publication. How we use AISend a correction

What happened
- Diffusion Controller attaches a lightweight network to a text-to-image model while the base model stays frozen, and Google likens it to a steering damper on a motorcycle.
- As the image forms, the add-on shifts the generation trajectory toward directions that raise a user-defined target, such as an artistic style or closer prompt alignment.
- A penalty guardrail is meant to stop that steering from distorting the image, so a prompted lizard gets its sunglasses without warped scales or proportions.
- The authors say the add-on can attach even to access-restricted, closed-source image models.
Compiled by The EngineerSomething wrong?How this is made
Why it matters
- decision Before counting on the add-on for a hosted model, a team has to find out what its provider exposes while an image is being generated.
- cost Adopting it means building or buying a reward that scores finished images for your own target, because the fine-tuning learns from nothing else.
- capability With the base weights frozen, a team can tune prompt adherence for its own use without retraining or modifying the base model.
The lizard is a fair test [14]. Ask for "a lizard wearing sunglasses" and a base model can drop the sunglasses, or distort the face to fit them in [14]. Hsu and Ryu, the Google Research engineers behind the September 29 post [1], wrote that the framework "strikes a balance" between the new preference and the base model's image quality [13]. Balance is the accurate word. The guardrail that protects the lizard's scales [10] also limits how hard the controller can push for the sunglasses. Adherence and quality still pull against each other. The difference is that the framework optimizes the trajectory shifts with feedback [5]. The authors say the older, fragmented tools left engineers to balance the two by guesswork [12].
Both fine-tuning methods learn from one reward score on the final image [6]. The controller will push toward whatever that score rewards, including its mistakes. For Google's win rate to carry over to another team's prompts, two things have to hold. That team's reward has to score adherence the way its own reviewers would. Its base model also has to behave like the baseline Google compared against [2]. The results summary does not name the industry standard the add-on beat [8].
The closed-model claim [7] needs the most checking. The authors note that changing a model's behaviour usually takes white-box access to its internal settings, and that the best image models are often corporate secrets they call black boxes [9]. Diffusion Controller acts during generation, adjusting the trajectory as the image forms [5]. A product sold on the secrecy of its internals is an odd place to attach something that steers them. If an endpoint returns only a finished picture, a controller that adjusts the trajectory mid-generation has no intermediate state to act on.
The part I'd expect to last is the framing. Classifier-free guidance changes the prompt's influence at inference time, while LoRA adapters, reward-weighted regression and policy gradients change the model through fine-tuning [11]. The authors wrote that these were treated as unrelated fixes, and that the field lacked "a single, principled mathematical language" for them [12]. Google presents Diffusion Controller as that language, recasting the whole denoising process as one continuous control problem [3]. If the formulation holds up, a guidance scale and a LoRA run can be compared in the same terms.
What to watch
- Add-on results measured on a named closed model, with a margin against a named baseline such as classifier-free guidance.
- Any hosted image API that exposes intermediate generation states or a steering hook an outside network can use.
- A paper or code release naming the reward and the judging setup behind the 90% win rate.