This blog post is based on a column that originally ran in The 74 on Feb. 5, 2026. In a landscape often dominated by sobering headlines about national test scores, a recent op-ed of math achievement on the National Assessment of Education Progress (NAEP) by L Burleigh, unearthed an underlying issue: the “math problem” in American classrooms isn’t actually about the numbers. It’s about a fundamental gap in how students are taught to think. According to Burleigh, U.S. students are currently facing a “comprehension crisis.” For decades, the education system has leaned heavily on procedural fluency, or teaching students to follow a set of memorized steps to reach an answer. While this might help a student pass a specific quiz, that knowledge isn’t correctly applied in the context of a problem that doesn’t fit the expected pattern. When the “recipe” fails, the student is left without the conceptual tools to navigate the kitchen. Engineering a Solution: The LEVI “Moonshot” In a landscape often dominated by sobering headlines about national test scores, a recent op-ed of math achievement on the National Assessment of Education Progress (NAEP) by L Burleigh, unearthed an underlying issue: the “math problem” in American classrooms isn’t actually about the numbers. It’s about a fundamental gap in how students are taught to think. According to Burleigh, U.S. students are currently facing a “comprehension crisis.” For decades, the education system has leaned heavily on procedural fluency, or teaching students to follow a set of memorized steps to reach an answer. While this might help a student pass a specific quiz, that knowledge isn’t correctly applied in the context of a problem that doesn’t fit the expected pattern. When the “recipe” fails, the student is left without the conceptual tools to navigate the kitchen. Learn more about the LEVI Math cohort are doubling math learning, from AI chatbots to digital avatars. From Drills to Dialogue: The Role of AI A standout example of this work is LEVI grantee Eedi. Traditional math software often functions as a high-tech version of a flashcard by marking an answer “correct” or “incorrect” and moving on. Eedi, however, uses advanced AI to diagnose the why behind a wrong answer. By identifying specific misconceptions, the platform acts less like a grader and more like a tutor. It asks guiding questions that force students to explain their reasoning, helping them build the mental scaffolding necessary to understand the “why” behind the “how.” This approach mimics high-dosage human tutoring at a scale that was previously impossible, directly addressing the comprehension gap identified in the original article. Building for the Long Term The takeaway is clear: if we want to move the needle on national math proficiency, we must stop treating math as a series of hurdles to jump over and start treating it as a language to be understood. By applying the rigors of engineering to the science of learning, LEVI and its teams are ensuring that students are building the critical thinking skills they will need for a lifetime.
Eedi Showing How AI Tutoring Can Deliver Personalized Learning Safely And Effectively
Education technology applications are continually being invented, and recent AI innovations have only increased that rate of innovation. While there are widely-accepted approaches to evaluate stable, late-stage products (e.g., Randomized Controlled Trials), there is much less clarity about how to conduct these evaluations at earlier stages of development. Given the potential risks in AI-powered solutions due to potential hallucinations and concerns about fairness, early evaluations are more important than ever.
AI Tutoring Can Safely And Effectively Support Students: An Exploratory RCT In UK Classrooms
AI Tutoring Can Safely And Effectively Support Students: An Exploratory RCT In UK Classrooms LearnLM Team, Google & Eedi One-to-one tutoring is widely considered the gold standard for personalized education, yet it remainsprohibitively expensive to scale. To evaluate whether generative AI might help expand access to thisresource, we conducted an exploratory randomized controlled trial (RCT) with 𝑵 = 165 students acrossfive UK secondary schools. We integrated LearnLM—a generative AI model fine-tuned for pedagogy—intochat-based tutoring sessions on the Eedi mathematics platform. In the RCT, expert tutors directlysupervised LearnLM, with the remit to revise each message it drafted until they would be satisfiedsending it themselves. LearnLM proved to be a reliable source of pedagogical instruction, with supervisingtutors approving 76.4% of its drafted messages making zero or minimal edits (i.e., changing only one ortwo characters). This translated into effective tutoring support: students guided by LearnLM performedat least as well as students chatting with human tutors on each learning outcome we measured. Infact, students who received support from LearnLM were 5.5 percentage points more likely to solvenovel problems on subsequent topics (with a success rate of 66.2%) than those who received tutoringfrom human tutors alone (rate of 60.7%). In interviews, tutors highlighted LearnLM’s strength atdrafting Socratic questions that encouraged deeper reflection from students, with multiple tutors evenreporting that they learned new pedagogical practices from the model. Overall, our results suggestthat pedagogically fine-tuned AI tutoring systems may play a promising role in delivering effective,individualized learning support at scale. Keywords learning, efficacy, safety, artificial intelligence, tutoring, randomized controlled trial
A Study To Evaluate The Effectiveness Of Eedi On Raising Attainment In Mathematics At KS3 (Year 7)
A Study To Evaluate The Effectiveness Of Eedi On Raising Attainment In Mathematics At KS3 (Year 7) Executive Summary This report evaluates the impact of Eedi, a digital mathematics platform, on raising maths attainment among Key Stage 3 students (Year 7). Eedi is an innovative EdTech platform offering over 60,000 diagnostic maths questions and engaging approximately 15,500 monthly users through its school and tutoring services. Enhanced by AI and video explanations, Eedi helps learners identify and address misconceptions while providing targeted additional support. Premium users also benefit from live, one-on-one tutoring, enabling students to request and receive assistance through the platform’s chat function. The impact evaluation included 20 schools, with randomisation conducted at the school level. The Eedi programme was implemented in intervention schools from Autumn 2023 to June 2024 to supplement classroom teaching, while control schools did not use the programme. Teachers assigned Eedi activities for completion either during lesson time or as additional homework tasks. The evaluation also incorporated a light-touch implementation and process review, drawing on online teacher surveys and Eedi usage data. Keywords large language models, intelligent tutoring systems, safety, system designGeneralization, Natural language processing, CollaborationMultiple Choice Question, Large Language Models, Humanin-the-loop.-
Improving The Validity Of Automatically Generated Feedback Via Reinforcement Learning
Improving The Validity Of Automatically Generated Feedback Via Reinforcement Learning Abstract Automatically generating feedback via large language models (LLMs) in intelligent tutoring systems and online learning platforms has the potential to improve the learning outcomes of many students. However, both feedback generation and evaluation are challenging: feedback content has to be valid especially in subjects like math, which requires models to understand the problem, the solution, and where the student’s error lies. Feedback also has to be pedagogically valid to reflect effective tutoring strategies, such as explaining possible misconceptions and encouraging the student, among other desirable features. In this work, we address both problems of automatically generating and evaluating feedback while considering both correctness and alignment. First, we propose a rubric for evaluating math feedback and show that GPT-4 is able to effectively use it to annotate human-written and LLM-generated feedback. Second, we propose a framework for feedback generation that optimizes both correctness and alignment using reinforcement learning (RL). Specifically, we use GPT-4’s annotations to create preferences over feedback pairs in an augmented dataset for training via direct preference optimization (DPO). We show that our methods significantly increase the correctness and alignment of generated feedback with Llama 2, an open-source LLM, qualitatively analyze our generation and evaluation systems using case studies, and outline several areas for future work Keywords Generalization, Natural language processing, CollaborationanalytiFeedback Generation, Human Preference Alignment, Math Education, Reinforcement Learning
Math Multiple Choice Question Generation Via Human-Large Language Model Collaboration
Math Multiple Choice Question Generation Via Human-Large Language Model Collaboration Abstract Multiple choice questions (MCQs) are a popular method for evaluating students’ knowledge due to their efficiency in administration and grading. Crafting high-quality math MCQs is a labor-intensive process that requires educators to formulate precise stems and plausible distractors. Recent advances in large language models (LLMs) have sparked interest in automating MCQ creation, but challenges persist in ensuring mathematical accuracy and addressing student errors. This paper introduces a prototype tool designed to facilitate collaboration between LLMs and educators for streamlining the math MCQ generation process. We conduct a pilot study involving math educators to investigate how the tool can help them simplify the process of crafting high-quality math MCQs. We found that while LLMs can generate well-formulated question stems, their ability to generate distractors that capture common student errors and misconceptions is limited. Nevertheless, a human-AI collaboration has the potential to enhance the efficiency and effectiveness of MCQ generation. Keywords Multiple Choice Question, Large Language Models, Human in-the-loop.Generalization, Natural language processing, CollaborationMultiple Choice Question, Large Language Models, Humanin-the-loop.-