ToolGrad: answer-first tool-use dataset generation with textual gradients
Most synthetic tool-use dataset generation starts with a user request and then searches for a way to answer it, wasting effort on failures. ToolGrad reverses the order: it assembles a working chain of API calls first, guided by textual gradients, and writes the matching request afterwards.

Why is tool-use dataset generation hard?
Teaching a language model to call tools needs many examples that pair a user request with the exact sequence of API calls that answers it. Human annotation does not scale, so most pipelines generate the pairs synthetically, and the usual way of doing that wastes much of its effort.
The common recipe, used by ToolBench and ToolACE among others, works query first. An LLM samples some APIs and imagines a request they might serve, then a search agent explores tool calls by trial and error, often with depth-first search, until it finds an answer or gives up.
The authors point to two costs. The exploration is expensive by design, because the useful trajectory has to be distilled from a large search, and nothing guarantees success, so failed searches burn agent calls without producing a sample. They also found that imagined requests can be unsolvable with the tools on offer, and that successful search traces sometimes contain wrong steps that then become training targets.
What does answer-first generation change?
It reverses the order. ToolGrad first builds a chain of API calls that has actually run successfully, then asks an LLM to write a user request and a final response that fit that chain. Describing a known solution is far easier than discovering one, and it takes a single LLM step.
The authors compare the inversion to summarising: turning a detailed workflow into the shorter, vaguer message a person might type. Because every sample starts from at least one tool call that worked, and failed steps never enter the data, each example is verified by construction.
The open question this creates is how to assemble useful multi-step chains directly from a large library of APIs without a request to steer by. That is the job of ToolGrad's iterative loop.
Borrowing textual gradients from prompt optimisation
ToolGrad adapts the idea of textual gradients from TextGrad, a framework for prompt optimisation. In TextGrad, an LLM critic writes plain-language feedback about a prediction, and another LLM edits the prompt in response, loosely mirroring gradient descent.
The authors stress that these are not mathematical gradients. In ToolGrad, the feedback signal is a choice rather than a critique: in each iteration a selector picks the one API worth adding, and that discrete pick plays the role of the gradient that moves the dataset sample forward.
The analogy has a limit the authors acknowledge. Optimising a dataset and the model trained on it at the same time would be a bi-level problem, so ToolGrad relies on LLM feedback to build the data without training a model inside the loop.
How does ToolGrad build a tool-use chain?
Each iteration runs four modules in sequence: propose, execute, select and update. Starting from the current sample, a triple of user query, API workflow and response, the loop adds at most one API per step and repeats for 10 iterations.
The result of a full run is one training sample: a natural-sounding request, a verified set of API chains with their inputs and outputs, and a final answer grounded in those outputs.
- API proposer. Reads a random mini-batch of 50 API descriptions and nominates up to 3 that could extend the workflow, each with an instruction for how to use it. It cannot call tools itself, which keeps this filtering step cheap.
- API executors. One tool-calling agent per proposal runs the API in parallel and returns a report with the full request history and a success flag. This is the most expensive step, which is why the proposer filters first. Calls time out after 10 seconds.
- API selector. Reviews the successful reports and picks the single most valuable API, then decides whether it extends an existing chain or starts a new one. Failed executions are ignored.
- Workflow updater. Appends the chosen API to the selected chain without an LLM, then uses an LLM to rewrite the user query and response so the sample stays coherent with the larger workflow.
From ToolGrad-500 to fine-tuned Gemma 3 models
The authors ran the loop 500 times with different seeds, using gemini-2.5-flash-lite for every LLM role because it is inexpensive and responds quickly. The output is ToolGrad-500, a set of 500 samples drawn from ToolBench's API library after two rounds of filtering left 15,368 usable APIs.
To mimic a realistic setting where a model sees more tools than it needs, each sample is padded with similar but unneeded APIs found by embedding similarity, giving 10 candidate tools. The authors also added 20% more negative examples in which none of the offered tools fit and the correct output is an empty list.
The data is formatted for single-turn tool use: given OpenAI-style tool definitions, the model must output every needed call at once as Python-style function calls, rather than one call per turn as ReAct agents do. Gemma 3 models at 1B, 4B and 12B parameters were then fine-tuned with supervised learning, the larger two through LoRA adapters.
Training was light. The authors report 1, 1.67 and 2.67 GPU hours on four A100 GPUs for the three ToolGrad models, compared with 29, 68 and 370 GPU hours on eight H100 GPUs for the matching models trained on ToolBench data.
How efficient is ToolGrad's data generation?
In the authors' comparison with ToolBench's depth-first search pipeline, ToolGrad produced a usable sample 99.8% of the time against 63.8%, while generating longer chains with fewer tool calls.
Chains averaged 3.4 ground-truth tool uses against 2.1 for depth-first search. LLM invocations were roughly level, 63.9 against 64.5, while tool calls fell from 34.3 to 20.0. The rare failures, 0.2% of runs, happened when none of the 3 selected APIs returned a successful response across all 10 iterations, leaving an empty sample.
The authors also tested runs of 4, 8, 12 and 16 iterations. Pass rate tended to level off between 8 and 12, which supports the choice of 10, but the share of unique tool uses exposed a weakness: the generator tends to produce similar tool uses across samples.
How do ToolGrad models perform on BFCL and ToolBench?
Fine-tuning on ToolGrad-500 improved every Gemma 3 size on both benchmarks, the authors report. On the Berkeley Function Calling Leaderboard (BFCL), whose tools barely overlap with ToolBench's, the 1B, 4B and 12B models gained 8.1, 8.0 and 6.3 points over their base models.
ToolGrad-12B scored 83.1 on BFCL's single-turn tracks, 0.1 points below gemini-2.5-pro at 83.2 and ahead of Claude 4.5 Opus at 82.8 and GPT-5 at 74.4, according to the authors. It also beat the specialised tool-use models ToolACE and Hammer-2.1-7B, and gemini-2.5-flash-lite, the model that generated its training data.
Gains were largest on BFCL's non-live, synthetic categories, at 14.19, 11.34 and 8.37 points, but live categories built from real user data also improved, by 1.93, 4.74 and 4.22. The authors single out the multiple parallel category, which combines choosing among tools with issuing several calls at once, as the clearest beneficiary.
On ToolBench's hardest cross-category track, scored by LLM judges, ToolGrad-12B and ToolGrad-4B ranked first and second at 19.6 and 17.6, above every proprietary model tested. A check with two human raters on 8 queries found that the averaged judge scores tracked human ratings closely.
The authors highlight a self-improvement effect: even the 1B model outscored its teacher on ToolBench, at 14.1 against 6.9. In their view, answer-first data lets a student exceed its teacher because training is not limited to problems the teacher could solve by search.
Limitations, the scaling plateau and open questions
The authors name four limitations. The fine-tuning data contains no reasoning steps, so the models are not suited to ReAct or depth-first search inference; only supervised fine-tuning was tested, not reinforcement learning; generated requests may not sound like real people; and performance stopped improving at a small dataset size.
The scaling result stands out. When the authors varied the training set for Gemma-3-4B from 100 to 2k samples, with steps at 500, 1k and 1.5k, every version beat the base model, but BFCL accuracy rose and then fell as the set grew. They point to one major reason, the lack of memory across runs: each sample is generated independently, so the framework keeps proposing similar tool uses.
Scope is also narrow. Evaluation covers BFCL's single-turn v1 and v2 tracks; multi-turn and agentic tracks are left for future work, as is post-processing generated queries for more natural, varied phrasing.
Our analysis: most of the BFCL gain sits in the non-live, synthetic categories, which resemble the generated training data, while gains on live categories built from real user requests are smaller. Teams adopting the recipe should measure results on their own traffic, including how often a fine-tuned model calls a tool when it should decline.
What it means for teams building AI agents
For organisations with their own internal APIs, the appeal is cost: a few hundred verified examples, generated by an inexpensive model and trained in a few GPU hours, lifted small open models close to large proprietary ones on a public benchmark. Answer-first generation also means every training example reflects calls that actually ran.
Our analysis: the method needs live, callable tools during generation, because every proposed API is executed. That suits sandboxed or read-only internal services, but tools with side effects, such as ones that send messages or change records, would need a safe test environment before a generator could explore them.
Together with the scaling plateau and the gap between synthetic and live gains, this suggests treating ToolGrad as a way to seed a tool-use dataset rather than as a complete pipeline, paired with real user requests and explicit tests for declining irrelevant tools.
Are the ToolGrad code, dataset and models available?
Yes. The authors have published the source code on GitHub under the Apache 2.0 licence, along with a Python package. The ToolGrad-500 dataset and the fine-tuned ToolGrad-1B, ToolGrad-4B and ToolGrad-12B models are on Hugging Face, where the model pages list the Gemma licence.
The dataset card lists 600 training sessions, 500 positive and 100 negative, plus 110 test sessions. The repository includes a demo on a Model Context Protocol (MCP) filesystem service that needs no GPU, scripts to reproduce the BFCL results and instructions for generating new data, which require a ToolBench API key.
The paper, by Zhongyi Zhou, Ruofei Du and six co-authors at Google, the University of Tokyo, RIKEN AIP and Tohoku University, appears in Findings of ACL 2026.
Questions and answers
What is a textual gradient in ToolGrad?
It is a feedback signal that comes from an LLM's judgement rather than from calculus. TextGrad introduced the idea for prompt optimisation, where written feedback from an LLM critic guides edits to a prompt. ToolGrad adapts it to data generation: in each iteration, an API selector reviews execution reports and picks the one API to add to the workflow. That discrete selection moves the sample forward, playing the part a numerical gradient plays when training a model.
How is ToolGrad different from ToolBench?
ToolBench generates a user request first and then searches for a tool-use solution with depth-first search, which can fail or produce flawed traces. ToolGrad builds a working chain of API calls first and writes the request afterwards. In the authors' comparison using ToolBench's own API library, ToolGrad reached a 99.8% pass rate against 63.8%, produced chains averaging 3.4 tool uses against 2.1, and needed fewer tool calls during generation.
Which models did the authors fine-tune with ToolGrad data?
They fine-tuned Gemma 3 models with 1B, 4B and 12B parameters using supervised learning on ToolGrad-500, which was generated with gemini-2.5-flash-lite. The 4B and 12B models were trained through LoRA adapters. Each model trained for three epochs, and the authors report total training costs of 1, 1.67 and 2.67 GPU hours on four A100 GPUs.
Can ToolGrad models handle multi-turn agent tasks?
Not as evaluated. The training data covers single-turn tool use, where the model emits every call at once, and contains no reasoning steps. The authors evaluated BFCL's single-turn v1 and v2 tracks and left multi-turn and agentic tracks for future work. They also note the models are limited when run inside ReAct or depth-first search frameworks, so multi-step agent use would need further data or training.
References
- Zhou, Z., Uehara, K., Zhang, H., Zhou, J., Gu, L., Du, R., Xu, Z., & Harada, T. (2026). ToolGrad: Efficient tool-use dataset generation with textual "gradients". Findings of the Association for Computational Linguistics: ACL 2026. arXiv:2508.04086. https://aclanthology.org/2026.findings-acl.950/ (external site)
- Zhou, Z. (2026). ToolGrad [Computer software]. GitHub. https://github.com/zhongyi-zhou/toolgrad (external site)
- Zhou, Z. (2026). ToolGrad-500 [Data set]. Hugging Face. https://huggingface.co/datasets/zhongyi-zhou/toolgrad-500 (external site)
- Yuksekgonul, M., Bianchi, F., Boen, J., Liu, S., Huang, Z., Guestrin, C., & Zou, J. (2024). TextGrad: Automatic "differentiation" via text. arXiv:2406.07496. https://arxiv.org/abs/2406.07496 (external site)
- Qin, Y., Liang, S., Ye, Y., Zhu, K., Yan, L., Lu, Y., Lin, Y., Cong, X., Tang, X., Qian, B., Zhao, S., Hong, L., Tian, R., Xie, R., Zhou, J., Gerstein, M., Li, D., Liu, Z., & Sun, M. (2023). ToolLLM: Facilitating large language models to master 16000+ real-world APIs. arXiv:2307.16789. https://arxiv.org/abs/2307.16789 (external site)
- Gorilla project, University of California, Berkeley. (n.d.). Berkeley Function Calling Leaderboard (BFCL). https://gorilla.cs.berkeley.edu/leaderboard.html (external site)
Original article
Zhou, Z., & Du, R. (2026, 10 September). ToolGrad: Efficient tool-use dataset generation with textual "gradients". Google Research Blog. https://research.google/blog/toolgrad-efficient-tool-use-dataset-generation-with-textual-gradients/ (external site)