A new educational-chatbot benchmark gives schools a way to compare answers before putting AI in front of students. The decision it leaves open is equally important: how much should a school value a correct first response, a successful correction, or resistance to unreliable behavior?
Researchers at the University of Namur published their AI Score in Scientific Reports on September 28. The peer-reviewed paper is available as an early version that may receive further edits. Its weighted measure combines initial performance, robustness, self-correction and lack of reliability. Tests used multiple-choice questions and course materials across first-year optics, law and history.
TENS Magazine’s analysis is that schools should read such a score as a profile of tradeoffs, then test whether those tradeoffs fit their teaching arrangements. A single total can support comparison, but the conditions under which a chatbot earns its points determine what that comparison means.
Keep the components beside the total
The Namur researchers evaluated six platforms using retrieval-augmented generation and class-specific resources. Their published abstract describes the framework as a pre-deployment assessment and identifies longer-term studies and qualitative assessment of answers as directions for further work. It does not establish that the highest-scoring chatbot produces the greatest student learning.
Consider a hypothetical comparison, not a result from the paper: one tutor answers correctly immediately but responds poorly when challenged; another starts less reliably but improves after correction. A combined score might favor either, depending on the weights. In a classroom with close teacher supervision, correction may be practical. During independent homework, the student may never recognize the need to challenge the response.
TENS proposes that a school publish the component results alongside its chosen weighting and explain who is expected to detect mistakes. This would make a hidden assumption visible: a system’s ability to recover is useful only when someone or something triggers the recovery. The score should not silently assume expert oversight where none is available.
A second proposed check is sensitivity to those weights. Before selecting a tool, a school could ask whether modest changes in priorities change the preferred candidate. If they do, the decision should be described as dependent on the teaching context. That would be more informative than treating a narrow aggregate lead as a universal recommendation.
The student and the chatbot take different tests
Two earlier experiments show why a chatbot assessment needs a separate learning assessment. A 2025 Scientific Reports trial by Harvard researchers Greg Kestin, Kelly Miller and colleagues compared a specially designed AI tutor with active classroom learning in college physics. Students using the tutor learned more in less time in that setting. The tutor incorporated pedagogical practices; the finding was not a test of every general-purpose chatbot.
A separate 2025 study in the Proceedings of the National Academy of Sciences, led by Hamsa Bastani and colleagues, examined high-school mathematics. Access to AI improved practice performance, yet students using the unrestricted version performed worse when the tool was removed. Teacher-informed safeguards largely mitigated that harm, without establishing a positive unassisted-exam effect for the guarded version.
These findings should not be compressed into a verdict that AI tutoring works or fails everywhere. The populations, subjects, designs and instructional settings differ. Together they sharpen the question a school must ask: does a tool help learners become more capable, or mainly help them complete the task while it remains available?
Test the assumption behind the purchase
TENS proposes connecting a benchmark profile to a small, explicitly defined classroom evaluation. The school would first identify the failure it most needs to avoid, such as an incorrect explanation surviving an entire homework session. It would then test the intended tutor configuration with the actual course resources and supervision arrangements, rather than carrying over a platform ranking without checking the setup.
The learning comparison could include an assessment after assistance is removed and a later exercise using unfamiliar examples. Those would address independence and retention separately. These are proposed checks, not outcomes demonstrated by the September paper or a TENS classroom experiment.
Teacher effort also belongs in the record. If one configuration needs frequent intervention to reach acceptable answers, its final accuracy can conceal a substantial support requirement. Logging the kinds of corrections and who supplies them would help distinguish a tutor that supports a lesson from one that shifts work onto teachers.
The useful next step is therefore a decision record that connects component scores, instructional priorities and observed student outcomes. A benchmark can make the first comparison more disciplined. Its strongest educational use would be to expose the assumptions that still need testing before a school treats an attractive number as evidence of a better tutor.
Image: Illustrative software photograph, not a tested tutoring platform. Markus Spiske via Wikimedia Commons. CC0 1.0 Universal Public Domain Dedication. Center-cropped from 5,760 × 3,840 pixels to 16:9 and resized to 2,400 × 1,350 pixels; no generative or content-altering edits.
