Learning interactives: how generative UI for education turns lesson goals into guided simulations
Learning interactives is a Google Research framework that applies generative UI for education, asking Gemini to build guided STEM simulations from a teacher's request. A loop of automated critics raises the share of usable simulations sharply, and teachers vet each one before it joins the public library.

Why is interactive practice so hard to produce?
Interactive simulations help students learn by doing, but they are costly to build, scarce and hard to adapt to a particular class. In a Google Research post (17 September 2026), Gal Elidan and Yael Haramaty describe research, detailed in an accompanying technical report, that tests whether generative models can close that gap without weakening the teaching.
The case for active learning is long established. The report draws on Dewey, Piaget, Bruner and Papert, and on the ICAP framework of Chi and Wylie (2014), which links interactive engagement with deeper understanding and longer retention than passive reading or listening. Koedinger and colleagues (2015) summed it up in a phrase the post borrows: "Learning is not a spectator sport."
The report also names two risks. Generative UI used straight out of the box tends to produce isolated widgets rather than structured lessons with several stages. And unguided discovery works poorly, as Kirschner, Sweller and Clark (2006) argued, while unrestricted use of AI assistants can tempt students to hand the thinking over to the tool, a tension the report, citing Jose and colleagues (2025), frames as enhancement versus erosion.
What is generative UI for education?
Generative UI means a model builds the interface itself, as working code, instead of filling a layout designed in advance. In education, that lets a teacher describe a topic and receive a bespoke simulation. Learning interactives is the authors' framework for doing this with Gemini under pedagogical guardrails, so that the result is a guided lesson rather than a one-off widget.
Earlier Google Research work by Leviathan and colleagues (2026) showed that a well-prompted model with the right tools can produce custom interfaces for almost any prompt. The education report argues that generic output falls short for complex lessons, which is why it wraps the model in structure built on four learning science principles, each mapped to a design choice.
- Curriculum alignment. Gemini drafts precise learning objectives from the teacher's request. The teacher can edit them, and they must be approved before the rest of the interactive is generated.
- Agency and motivation. Each topic is split into levels with clear goals that grow harder step by step, borrowing from game-based learning to balance challenge with a sense of progress.
- Guidance and scaffolding. An introduction activates prior knowledge, a toolbox holds the relevant formulas and theory, and hints come in tiers, so students get help without being handed the answer.
- Formative feedback. The simulation responds instantly to each action, explanatory feedback ties outcomes to the underlying concept, and debriefs and worked solutions follow the student's own attempt.
How does the learning interactives pipeline work?
Generation runs in four stages, and the first two produce text only. A simulation is planned fully in words, with its levelled goals checked, before any interface code is written. That way the most error-prone step, writing the interface, starts from a validated specification.
In the report's electrical circuits example, a high school request leads to objectives on RC time constants, LC resonance, p-n junctions and photonics, and to a design in which the student tunes each stage of a signal path on a microchip so that it carries a target frequency without distortion.
- Simulation idea. From a request naming the subject, grade, topic and any extra requirements, Gemini writes learning objectives and then a detailed design listing the entities, how they relate, the graphs to show, the visual features and every knob and button.
- Levelled goals. Gemini writes five goals one at a time, each conditioned on those before it. A critique checks that every goal is concrete, clear, harder than the last and not too big a jump, and failed goals are regenerated. The report says this stage costs little run time.
- Interface generation. Gemini writes the simulation in HTML and JavaScript. Critics then review both the code and the page as rendered in Chrome, and the model revises it over repeated rounds.
- Guidance. Finally, the system writes an onboarding tour, hints and a solution for each level, and feedback for both successful and incorrect attempts.
Generate, critique, repair: the refinement loop
The authors report that, even with a validated idea and goals, a single prompt produced a learning interactive meeting their pedagogical requirements only 3.5% of the time. Their remedy is a generate-then-refine loop: critics test each draft, the model repairs whatever they flag, and the cycle repeats.
The critics fall into four groups: visual, solution, telemetry and mechanical. The post gives a flavour of the questions they ask. Do the levels cover the objectives and get progressively harder? Do the buttons work, and can each level actually be solved? Is anything on screen redundant or distracting? Some checks act as agents: a solvability critic opens Chrome and plays the simulation as a student would, including adversarial moves such as pushing knobs to their limits.
With this loop, the share of simulations passing every critique rose from 3.5% after the first generation to 69.3% after 10 rounds, according to the report. Progress within individual critique groups was not steady, because fixing one part of an interactive can upset another. The post is open about the cost: the loop makes generation slower in exchange for closer adherence to the quality criteria.
What did teachers think of the learning interactives?
In the authors' usability study, 12 US teachers rated every simulation they received at 7 or above for usability, with an average above 8.1, according to the report. The post summarises the same study as an average quality rating of 8 out of 10.
The group comprised six men and six women: four mathematics teachers, two each in biology, physics and environmental science, and one each in astronomy and chemistry. Each teacher submitted three requests, 36 in all, and approved most of the generated learning objectives unchanged. For each request the team generated five candidates and internal raters picked the best one; for 3 of the 36 requests, no candidate was judged good enough to send.
Teachers valued being able to tailor a simulation to their own class. One high school science teacher, quoted in the post, said: "I've never been able to differentiate any of the simulations because it's just, you get what you get". Others said the tiered hints and worked solutions resembled the one-to-one help they give students themselves. Requests for improvement included alignment with learning standards and feedback on incorrect attempts, which the report says has since been added.
How did expert raters score the simulations?
For a larger check that the authors describe as unbiased, pedagogical experts chose 40 STEM requests ranging from upper primary to undergraduate level. The team ran five generation attempts per request, and every interactive that passed all critiques went to two raters who teach the relevant subject. The authors report an acceptance rate of 86%, with most accepted simulations rated good or excellent.
Raters first decided whether to accept a simulation, then scored it against rubrics covering the learning objectives (factuality, curriculum alignment, whether they meet SMART criteria, clarity and progression) and the user and learning experience (alignment, cognitive load, metacognition, usability and accessibility, and aesthetics). Scores ran from fail (0) through pass (1) and good (2) to excellent (3). The post adds that the raters were STEM teachers in the UK.
Results varied by subject. Physics and chemistry came closest to excellent on average. Biology scored lower, especially on aesthetics and user experience, because many biology topics, such as the stages of the cell cycle, depend more on remembering details than on manipulating equations, so the simulations drifted toward a textbook sequence.
Limitations and open questions
The evidence so far concerns teachers' judgements, not student outcomes. The authors call the work a first step and say the most notable gap is evaluation with real teachers in real schools, including whether students actually learn more.
Analysis: the report describes the pipeline at a high level. It does not publish the prompts, the critic implementations or a specific Gemini model version, so the 3.5% and 69.3% figures cannot be reproduced independently from the sources alone.
- Small samples. The usability study involved 12 teachers in the US, and the expert rating covered 40 requests with two raters each.
- Imperfect generation. After 10 refinement rounds, 69.3% of simulations passed every critique, so a sizeable share still did not, and the loop lengthens generation time.
- Uneven subjects. Topics that rely on memorisation, as much of biology does, translate less naturally into simulations a student can manipulate.
- Scope. The public library holds over 30 interactives in English for STEM subjects, mainly for middle and high school, and standards alignment was still a requested feature in the study.
What can teams building generated interfaces learn from it?
The transferable lesson lies in the method rather than the subject. When a model has to generate working interfaces, the report's results suggest that the structure around the model matters as much as the model: plan in text first, check intermediate outputs against explicit criteria, test the rendered result by using it, and keep a qualified person as the final approver.
Analysis: these patterns apply to generated software in many settings, from internal tools to training material and customer-facing explainers. The study also shows why single-pass success rates mislead. The jump from 3.5% to 69.3% came from the evaluation loop, not from a different model. Teams taking a similar approach should budget for longer generation, measure what reaches users rather than what the model first produces, and confirm outcomes with the people the tool is meant to serve.
Library, pilot programme and availability
Google Research has released a public library of over 30 learning interactives in English, across topics in physics, chemistry, biology, mathematics, earth science and computer science, all generated by AI and vetted by teachers. Examples named in the post include Kepler's laws of planetary motion, data visualisation and projectile motion.
Schools using Google Workspace for Education can sign up for the Google for Education Pilot Program, through which teachers can request custom STEM interactives for their curriculum, learning goals and grade level. New interactives go back to the requesting teacher for review and join the library only after approval. The team also plans user research and field studies on learning gains and engagement. The sources do not mention a code release.
Questions and answers
What are learning interactives?
Learning interactives are AI-generated, browser-based STEM simulations built around a teacher's learning objectives. Each one splits a topic into levels of increasing difficulty and adds scaffolding: an introduction, a toolbox of formulas and theory, tiered hints, feedback on attempts and worked solutions. Google Research builds them with Gemini through a multi-stage generative UI pipeline, and teachers approve the objectives and vet the finished interactives before they appear in the public library.
How does generative UI differ from a chatbot answer?
A chatbot usually returns text inside a fixed interface. With generative UI, the model writes the interface itself, as working code with controls, visuals and behaviour suited to the request. Earlier Google Research work showed that well-prompted models can do this for a wide range of prompts. The learning interactives report adds pedagogical structure and automated critiques because generic output often falls short for lessons with several stages.
How reliable are AI-generated learning simulations?
The authors measured quality rather than asserting it. Automated critics check for visual, solution, telemetry and mechanical problems, and in the report 69.3% of simulations passed every critique after 10 refinement rounds, up from 3.5% after one. Expert teacher raters then scored factuality and other criteria and accepted 86% of the interactives they reviewed. Teachers vet every interactive before it enters the public library.
Do learning interactives improve student learning?
That has not been shown yet. The published evaluations measure teachers' judgements: usability ratings from 12 US teachers and rubric scores from expert raters. The authors state that evaluation in real schools, including how well students learn, is the most notable remaining need, and they plan field studies and a pilot through the Google for Education Pilot Program to measure learning gains and engagement.
Can teachers use learning interactives today?
Teachers can open the public library of over 30 teacher-vetted interactives in English, mostly for middle and high school STEM topics. Custom requests currently go through the Google for Education Pilot Program, which schools using Google Workspace for Education can join. Teachers in the pilot request interactives for their own curriculum and review each one before it is added to the library.
References
- Kovshov, A., Choudhury, A., Iurchenko, A., Keeling, A., Hassidim, A., Shasha Evron, A., Çakmakli, A., Akrong, D., Olanubi, F., Elidan, G., Li, I., Lerer, I., Chou, K., Hackmon, L., Gordon, M., Kerem, N., Efron, N., Singh, P., Levitt, R., . . . Lev, Y. (2026). Harnessing generative UI for education: Tailored learning interactives. arXiv:2609.20738. https://arxiv.org/abs/2609.20738 (external site)
- Leviathan, Y., Valevski, D., Kalman, M., Lumen, D., Segalis, E., Molad, E., Pasternak, S., Natchu, V., Nygaard, V., Venkatachary, S., Manyika, J., & Matias, Y. (2026). Generative UI: LLMs are effective UI generators. arXiv:2604.09577. https://arxiv.org/abs/2604.09577 (external site)
- Chi, M. T. H., & Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4), 219-243. https://doi.org/10.1080/00461520.2014.965823 (external site)
- Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work: An analysis of the failure of constructivist, discovery, problem-based, experiential, and inquiry-based teaching. Educational Psychologist, 41(2), 75-86. https://doi.org/10.1207/s15326985ep4102_1 (external site)
- Koedinger, K. R., Kim, J., Jia, J. Z., McLaughlin, E. A., & Bier, N. L. (2015). Learning is not a spectator sport: Doing is better than watching for learning from a MOOC. In Proceedings of the Second (2015) ACM Conference on Learning @ Scale (pp. 111-120). ACM. https://doi.org/10.1145/2724660.2724681 (external site)
- Jose, B., Cherian, J., Verghis, A. M., Varghise, S. M., S, M., & Joseph, S. (2025). The cognitive paradox of AI in education: Between enhancement and erosion. Frontiers in Psychology, 16, 1550621. https://doi.org/10.3389/fpsyg.2025.1550621 (external site)
Original article
Elidan, G., & Haramaty, Y. (2026, 17 September). The future of practice: Enabling teachers to create learning interactives with generative UI. Google Research Blog. https://research.google/blog/the-future-of-practice-enabling-teachers-to-create-learning-interactives-with-generative-ui/ (external site)