Skip to content
AI

AI Classroom Simulations Need Evidence Beyond Working Software

Google’s learning interactives bring separate tests for software quality, teacher review effort and student understanding into focus.

Share Email
Colorful software code displayed on a computer monitor
Source code displayed on a computer monitor. Photo: Markus Spiske via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped from 5,760 × 3,840 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative or content-altering edits.

AI can now build a classroom simulation around a teacher’s request. The harder question is what the resulting activity asks a student to think through. Google Research’s September 17 release of learning interactives brings that distinction into focus: generating working educational software and demonstrating that students learn from it require different evidence.

Google has released an English-language library of more than 30 teacher-reviewed STEM interactives, aimed at middle and high school. It also announced a school pilot through Google for Education. The experiment uses generative user interfaces to turn instructional objectives into guided simulations, with teachers approving objectives and reviewing the resulting material.

TENS Magazine’s analysis reads this development against the ICAP framework developed by Arizona State University researchers Michelene Chi and Ruth Wylie. The central test is whether customization improves the intellectual work a lesson demands, while keeping the cost of checking that lesson manageable. A larger supply of activities would matter, but activity counts alone cannot answer either question.

A working activity is the first threshold

Google’s September 16 technical report describes repeated automated critique and correction of generated interfaces. In the reported generation evaluation, the share passing all critiques rose from 3.5 percent after initial generation to 69.3 percent after ten improvement rounds. Those checks address the generated software; they are not measurements of student learning.

The practical implication is that educational generation should be treated as a production process with rejected outputs. A teacher deciding whether to rely on it needs to know how much review remains after the software says an activity is ready. The time spent requesting, checking and revising a simulation belongs in the same comparison as the time saved building it.

Consider a hypothetical lesson about changing a circuit’s resistance. A functioning slider and a correctly moving graph would establish that the interface responds. A useful classroom review would additionally ask whether the displayed relationship is correct, whether the task exposes the intended concept, and whether an unexpected student action produces misleading feedback. These are proposed evaluation questions, not observations from a TENS classroom trial.

Interaction needs an intellectual purpose

Chi and Wylie’s 2014 ICAP paper distinguishes passive reception, active manipulation, constructive generation and interactive dialogue. Its categories concern how learners engage with material. Merely operating a computer interface does not automatically place a lesson in the framework’s most demanding category.

That distinction changes how an AI-generated activity should be judged. In the circuit example, moving a control until a success message appears might demonstrate successful exploration without showing why the answer works. Asking the student to predict the result first, explain a mismatch, and defend a revised explanation would create more revealing evidence for the teacher.

The same screen could support either lesson. The educational difference would lie partly in the prompt, sequence and follow-up discussion. This is why TENS’s reading places the teacher’s instructional objective above the visual novelty of the generated interface. A polished animation is useful only insofar as it helps students do the intended thinking.

Google’s design includes progressively challenging tasks and scaffolded guidance. Those features make the research relevant to this distinction, but their presence does not establish how individual students use them. A pilot could examine whether learners use hints to revise an explanation or simply to reach the next completion marker.

Teacher approval and learning evidence answer different questions

The technical report’s US usability study involved 12 teachers. For each request, internal raters selected among five generated candidates; three of 36 requests did not produce material considered good enough to send to the teacher. That selection process matters when interpreting positive responses: the feedback describes screened outputs rather than every initial generation.

For a school, a sensible next comparison would separate three outcomes: usable activities delivered per request, teacher time required per accepted activity, and what students can explain afterward. Improvement in one should not be reported as proof of improvement in all three. Strong filtering could produce attractive final materials while leaving the production burden unresolved.

Google says classroom studies of engagement and learning gains are still planned. Teacher judgments are valuable evidence about suitability and usability; student assessments would address a further claim. Neither should be dismissed, and neither should substitute for the other.

A useful pilot could therefore compare an AI-tailored activity with the material a teacher would otherwise use, under comparable lesson conditions. It could ask students to apply the concept to a new problem after the simulation is removed. That would help distinguish learning that travels beyond the interface from performance supported by its immediate cues.

The release opens a concrete opportunity to test whether teachers can obtain better-matched practice without becoming full-time software reviewers. Its educational value will become clearer when the evidence connects generation, review effort and independent student understanding. Until then, the library is a research resource and a starting point for classroom evaluation.

Image: Illustrative software photograph, not Google’s learning interface. Markus Spiske, via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped to 16:9 and resized to 2,400 × 1,350 pixels; no generative or content-altering edits.