Retrieve-for-Train: how a small diffusion model speeds up query fan-out retrieval
Retrieve-for-Train (R4T) moves the costly reasoning behind query fan-out retrieval out of the live search path. Reinforcement learning discovers good fan-out behaviour offline, and a 53.9M-parameter diffusion model then reproduces it in one pass, which the authors report is 12 to 20 times faster.

What is query fan-out retrieval?
Query fan-out retrieval splits one broad request into several narrower sub-queries, runs each one against a database and merges what comes back into a single set. It suits searches where people want a spread of options rather than one best match: someone searching for camping gear should see a tent, a sleeping bag, a stove and a headlamp, not ten versions of the same tent.
The paper calls this set-valued retrieval. Quality belongs to the collection as a whole: how diverse it is, how fully it covers the intent, how well its parts complement each other and whether every item really exists in the database. Such properties are non-decomposable, so scoring each result in isolation, as standard learning-to-rank does, cannot measure them.
Many different sets can satisfy the same broad intent, so there is rarely one correct answer to label. Conventional training data built from pairs of a query and its single best item is therefore a poor fit, and collecting labelled examples of good sets is costly and subjective.
Why do general language models struggle with query fan-out?
General-purpose language models can write sub-queries, but the authors identify two weaknesses. The sub-queries tend to repeat each other, and writing good ones takes long chains of reasoning tokens, adding latency that a live search box cannot absorb.
Grounding is a related problem. A generic model does not know the shape of a particular catalogue, so it can propose sub-queries that match nothing real in it.
- Paraphrastic collapse. Without knowledge of what a database holds, a model asked to expand "Bohemian festival style" may return little more than rewordings of it. The results come back homogeneous and miss facets such as suede boots, crochet dresses or fringe jackets that an expert stylist would add.
- Sequential latency. Autoregressive models emit one token at a time, and planning a good decomposition can take hundreds of intermediate reasoning tokens before the first search term appears. The cost compounds when many sub-queries are needed at once, which the authors say leaves a latency floor that conflicts with sub-second search.
How does Retrieve-for-Train work?
Retrieve-for-Train runs reinforcement learning once, offline, and turns what it learns into training data for a fast model. The authors describe RL here as an objective transducer: a way to convert a goal that is hard to label into ordinary supervised examples.
The diffusion model uses a variance-exploding formulation within the EDM framework, receives the query embedding through cross-attention and applies classifier-free guidance. For open-ended tasks its targets are the embeddings of retrieved items; for the weakly supervised task they are the embeddings of the learned sub-queries, so the model absorbs the decomposition strategy itself.
The authors test two ways to deploy the result. R4T-FOLM serves the RL-tuned language model directly, showing the quality the learned behaviour can reach while keeping autoregressive latency. R4T-Diffusion serves the distilled diffusion model, testing whether that behaviour survives in a single fast pass.
- Stage 1: train a fan-out language model. An open 4B-parameter model, Gemma3-4B or Qwen3-4B in the experiments, learns through reinforcement learning to write 10 sub-queries for each broad query. A frozen retriever runs them, and a set-level reward scores the combined result.
- Stage 2: synthesise supervision. The trained model is frozen and generates fan-outs for a large pool of queries, 128 samples per query at temperature 0.9. Each query is stored with a target set of embeddings, and no human labelling is involved. The target rows are shuffled during training because the order of a set carries no meaning.
- Stage 3: train a diffusion retriever. A 53.9M-parameter diffusion transformer learns to map a query embedding straight to a whole set of target embeddings. At query time it generates every retrieval direction together in one pass, and each output embedding is matched to its nearest item in the database.
Designing a reward for the whole result set
For open-ended queries with no ground truth, the reward is a weighted sum of three scores computed over the full set of sub-queries. Each term closes a shortcut the others leave open, which is why the authors describe them as mutual counter-anchors.
By default the weights are 0.6 for groundedness and 0.2 each for diversity and alignment. In the weakly supervised task, where each query comes with one plausible reference set, the reward is simply the share of reference items recovered by the union of fan-out results.
The fan-out model is optimised with group relative policy optimisation (GRPO), which scores each sampled output against others drawn for the same query, combined with soft PPO regularisation that adds forward and reverse KL penalties to keep open-ended generation stable.
- Groundedness. Rewards sub-queries whose embeddings sit close to a real item in the database, using the distance to the nearest neighbour, so the model is pushed towards things that exist.
- Diversity. Applies the Vendi Score, a diversity metric built on pairwise similarity, to one representative item retrieved per sub-query. A higher score means the set spreads across more distinct meanings.
- Alignment. Averages the cosine similarity between each sub-query and the original query, which stops the set drifting away from what the person actually asked for.

What happens when the reward leaves out diversity?
Training goes wrong quickly. In the authors' ablation, a model rewarded for groundedness alone learned to emit nonsense strings, such as a repeated phrase about line endings, that happen to sit near a database item. Adding alignment without diversity made it collapse into paraphrases of the original query even faster.
Only the full three-part reward produced stable training. Weighting matters as well: when groundedness dominated, diversity rose but alignment fell, and when alignment and diversity dominated, exploration was held back. A balanced mix let all three scores converge.
Our analysis: the training curves carry a warning for anyone tuning retrieval with reinforcement learning. In the paper's chart, the two hacked runs climb fastest and settle at the highest reward values, while the healthy run rises slowly. A rising reward curve on its own therefore says little about whether the learned behaviour is useful.

How well does Retrieve-for-Train perform?
The authors report that R4T outperformed every fan-out baseline within each model family on open-ended retrieval, and improved the balance between coverage and diversity on the weakly supervised task. The diffusion variant kept most of the quality of the RL-tuned language model.
Experiments use two datasets. Polyvore is a published fashion benchmark of user-curated outfits, searched with a CLIP-based image and text encoder over 21,888 collections for open-ended retrieval and 142,472 individual items for the weakly supervised task. The second is a proprietary set of expert-made music playlists, searched with the MuLan music and text embedding model over 8,522 playlist embeddings.
Baselines include retrieval with no fan-out, zero-shot fan-out from Gemini-2.5-Flash, Gemma3-4B and Qwen3-4B, and Best-of-N, which runs zero-shot fan-out several times (N=5) and keeps the output with the highest training reward.
For open-ended retrieval, an LLM judge rated diversity, alignment and groundedness on a 5-point scale. On Polyvore with Gemma3-4B, the average score rose from 38.5 for zero-shot fan-out and 40.9 for Best-of-N to 49.1 for R4T-FOLM, and diversity climbed from 56.0 to 76.8. On the music data, the Gemma-based average moved from 48.1 for zero-shot fan-out to 58.1 for R4T-FOLM.
For weakly supervised compositional retrieval, each Polyvore outfit became a reference set paired with an LLM-written broad query. R4T-FOLM on Qwen3-4B reached a Recall@5K of 20.9 and a Hit@5K of 64.6, against 15.7 and 52.1 for zero-shot Gemini-2.5-Flash. The diffusion variants gave up some recall for higher Vendi Scores, which the authors read as broader coverage of valid alternatives rather than lower quality.
How much faster is the diffusion retriever?
In the authors' benchmark, the diffusion retriever ran 12 to 20 times faster than autoregressive fan-out when producing 10 retrieval directions. For a batch of 8 queries it took 0.07 seconds against about 1.46 seconds; for a batch of 1,024 it took 4.21 seconds against nearly 50 seconds.
The language model's time is dominated by fixed overhead at small batches and then grows roughly in line with batch size. The diffusion model avoids token-by-token decoding because it denoises all target embeddings together in a continuous space, and its small size keeps serving memory low.
Best-of-N, the strongest baseline, moves the other way. It needs several independent fan-out runs for every query, which the authors say raises inference cost by an order of magnitude.

Limitations and open questions
The authors list four limitations: the offline RL stage is costly, the method needs goals that can be written as explicit rewards, open-ended quality was judged by a language model, and results may depend on the chosen models, embedding spaces and diffusion architecture.
On cost, the RL stage interacts with the retriever many times, and the authors note that this upfront work may be substantial for very large databases or ones that change often. On rewards, qualities such as creativity, novelty or cultural sensitivity may be hard to express as a single score; they suggest learning rewards from user feedback.
The paper also names an ethical risk. A carelessly specified reward could encode or amplify bias, and synthetic training data could spread that bias at scale, so the authors call for domain-specific bias audits and human oversight.
Our analysis: two details limit independent checking. The music results rely on a proprietary dataset, and the paper names Gemini-2.5-Pro as the judge in its main text but Gemini-2.5-Flash in an appendix.
Why it matters for enterprise search and recommendation
The core pattern travels well: spend compute once, offline, to discover behaviour that is too slow to compute on every request, then distil it into a small model that serves live traffic. For catalogue search, content discovery and recommendation slates, this could mean diverse result sets without a reasoning model in the request path.
R4T also addresses a common data gap. Many organisations hold a catalogue and embeddings but no labelled examples of what a good set of results looks like. The method replaces those labels with an explicit reward, which moves the effort into defining that reward carefully and auditing what it encourages.
Our analysis: teams weighing the approach would want to know how often the fan-out model must be retrained as a catalogue changes, how reward weights transfer to their own domain, and how to measure set quality without relying on an LLM judge alone.
Is the Retrieve-for-Train code available?
Neither the post nor the paper announces a release of code, trained models or the synthetic training data. The paper, by Pengcheng Jiang, Judith Yue Li and nine co-authors, is on arXiv under a CC BY 4.0 licence, and the authors' post describes it as an ICML 2026 paper.
The building blocks are public: Gemma3-4B and Qwen3-4B are open models, Polyvore is a published benchmark, and the paper's appendix lists training settings, including a 6-layer diffusion transformer with 16 attention heads and 256 sampling steps.
Questions and answers
What is the difference between R4T-FOLM and R4T-Diffusion?
Both variants share the same reinforcement learning stage. R4T-FOLM serves the RL-tuned fan-out language model directly, writing sub-queries token by token and calling the retriever for each, which gave the best measured quality but keeps autoregressive latency. R4T-Diffusion serves a 53.9M-parameter diffusion model trained on that model's outputs and generates all retrieval embeddings in one pass. The authors report it kept most of the quality while running 12 to 20 times faster.
Does Retrieve-for-Train need human-labelled training data?
No. The training pairs for the diffusion retriever come from the RL-tuned fan-out model, which generates sub-queries and retrieval targets for each broad query offline. Human judgement enters through the design of the reward, which encodes groundedness, diversity and alignment, and through the choice of its weights. In the weakly supervised task, existing curated sets such as Polyvore outfits act as reference sets for a coverage reward.
What is the Vendi Score?
The Vendi Score is a diversity metric for machine learning introduced by Friedman and Dieng in 2022. It measures the semantic breadth of a set from the similarities between its members' embeddings, so a higher score means the set covers more distinct content. R4T uses it twice: as the diversity term in the training reward and as an evaluation metric for weakly supervised retrieval across five independent runs.
Can query fan-out retrieval run in real time?
The authors' timings suggest the diffusion approach brings it within reach. For 10 retrieval directions, their diffusion model took 0.07 seconds for a batch of 8 queries and 4.21 seconds for a batch of 1,024, while autoregressive fan-out needed about 1.46 seconds and nearly 50 seconds. Production latency would also depend on the vector index, ranking and serving stack, which the benchmark does not cover.
References
- Jiang, P., Li, J. Y., Ryu, M., Hu, R. L., Su, K., Wan, Z. Y., Hebert, L., Peng, H., Han, J., Kuzmin, D., & Boutilier, C. (2026). Efficient, property-aligned fan-out retrieval via RL-compiled diffusion. arXiv:2603.06397. https://arxiv.org/abs/2603.06397 (external site)
- Friedman, D., & Dieng, A. B. (2022). The Vendi Score: A diversity evaluation metric for machine learning. arXiv:2210.02410. https://arxiv.org/abs/2210.02410 (external site)
- Han, X., Wu, Z., Jiang, Y.-G., & Davis, L. S. (2017). Learning fashion compatibility with bidirectional LSTMs. Proceedings of the 25th ACM International Conference on Multimedia, 1078-1086. arXiv:1707.05691. https://arxiv.org/abs/1707.05691 (external site)
- Huang, Q., Jansen, A., Lee, J., Ganti, R., Li, J. Y., & Ellis, D. P. W. (2022). MuLan: A joint embedding of music audio and natural language. arXiv:2208.12415. https://arxiv.org/abs/2208.12415 (external site)
Original article
Jiang, P., & Li, J. Y. (2026, 15 September). Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train. Google Research Blog. https://research.google/blog/bypassing-inference-bottlenecks-accelerating-complex-ai-search-with-retrieve-for-train/ (external site)