In a randomized trial of nearly 1,000 high school students in Turkey, students given an unguarded ChatGPT-style math tutor scored 17% worse on an unassisted exam than students who had never touched AI. (Source: PNAS, 2025) A trial of 2,379 undergraduates published later found that course-integrated AI tutor access cut final grades by 0.37 standard deviations. (Source: Annenberg Institute at Brown University, 2026) Both results are real, both are large, and both matter directly to DepEd and the private school networks preparing to scale AI tutoring in the Philippines.

The result everyone quotes
The study most often cited in favour of AI tutors came from an introductory physics course at Harvard. Students assigned to an AI tutor showed median learning gains more than double those of classmates in an active learning classroom, with an effect size between 0.73 and 1.3 standard deviations depending on the estimator used. (Source: Scientific Reports, 2025) Median time on task was 49 minutes against a 60-minute class, and 83% of students rated the tutor's explanations as equal to or better than their instructor's.
The design details matter more than the headline. This was not a chat window bolted onto a syllabus. Researchers engineered the system prompt for active learning, cognitive load management, and growth mindset, then discovered that prompt alone was insufficient because the model could not reliably scaffold multi-part problems. They embedded expert-written step-by-step solutions into the prompts and built a platform that walked students sequentially through each part of each problem, mirroring the instructor's in-class sequence.
What the negative results measured
The Turkish field experiment deployed two variants across roughly 50 classes. Both improved practice performance dramatically: the standard ChatGPT-style tutor by 48% and the safeguarded tutor by 127% over a textbook-only control. (Source: PNAS, 2025) The divergence appeared only when the laptops were closed.
On the closed-book exam, the unguarded group scored 17% below control. The safeguarded group, whose prompt was instructed to give hints and never the answer, landed statistically indistinguishable from students who had never used AI. (Source: PNAS, 2025) The unguarded model also produced a correct answer only 51% of the time on the practice problems, with logical errors accounting for 42% of failures. Interaction logs showed students asking "What is the answer?" and copying, while the safeguarded group asked for help and attempted answers independently.
The American university trial failed differently. Tutor access reduced learning management system participation by 0.90 standard deviations, meaning students stopped showing up to coursework at all, and estimated academic losses were larger for first-generation students. (Source: Annenberg Institute at Brown University, 2026) The authors described the mechanism as displacement, where the tutor absorbed the engagement that other learning activities would otherwise have captured.
Why the trials disagree
All three studies randomized something, but they did not randomize the same thing. The Harvard trial held pedagogy constant and varied the delivery medium, with expert content built in. (Source: Scientific Reports, 2025) The Turkish trial varied guardrails on a fixed underlying model. (Source: PNAS, 2025) The university trial varied access to a course-integrated tool in normal operating conditions. (Source: Annenberg Institute at Brown University, 2026)
None of them randomized an off-the-shelf chat assistant handed to unmotivated students, which is the most common deployment shape in schools today.
The Philippine constraint
Connectivity decides how much of this is even reachable. The 2024 National ICT Household Survey found that 48.8% of households had an internet connection, while two in every three individuals aged 10 and above used the internet somewhere. (Source: Philippine Statistics Authority, 2025) Roughly half of Filipino households cannot support an always-on tutor at home, and access is unevenly distributed below that national average.
This interacts badly with the negative findings. The subgroup that lost most in the university trial was first-generation students, the same group least likely to have reliable connectivity and a quiet place to study. (Source: Annenberg Institute at Brown University, 2026) A deployment that assumes a device, a connection, and slack in a student's schedule is being designed for a segment of the population, not for the system.
What the design evidence supports
Three interventions cost almost nothing and are backed by the trial designs themselves.
Guardrails beat access. The safeguarded tutor erased the exam penalty entirely relative to control at no measured cost to the outcome. (Source: PNAS, 2025) Prompting the model to hint rather than answer is a configuration change, not a procurement decision.
Participation is an early warning. A 0.90 standard deviation drop in learning management system log-ins is visible long before grades move. (Source: Annenberg Institute at Brown University, 2026) If platform activity falls after a tool launches, that is the signal, not a rounding error.
Teacher-authored grounding carries the gains. Both positive studies used problem-specific prompts written by subject experts rather than generic ones. (Source: Scientific Reports, 2025) The pre-instruction groundwork is not prompt polish, it is a teacher writing out the correct solution and the common wrong turns.
FAQ
Q: Does this mean AI tutors do not work?
A: The evidence says they work when built as a designed pedagogy and fail or actively harm when deployed as unrestricted access. A meta-review of 50 studies of earlier intelligent tutoring systems found these systems can match human tutoring when the instructional design is sound. (Source: Brookings Institution, 2026)
Q: Which should a Philippine school pilot first?
A: Practice sessions on topics already covered, with a hint-only tutor and a closed-book check the following day. That sequence makes the learning effect measurable within a single term instead of waiting for grade outcomes.
Q: What about students who already use ChatGPT at home without permission?
A: That is exactly the unguarded condition the Turkish trial measured, and the same 17% exam penalty applied. (Source: PNAS, 2025) The uncontrolled use already happening outside school is the version most likely to be costing learning.
Key Takeaway
The three trials do not disagree about whether AI can teach. They disagree about what a school is actually deploying when it turns one on. The variable that decides the outcome is not model capability, it is whether pedagogy, guardrails, and a teacher's authored content sit between the student and the chat box.
Before a single device is bought: which of the three interventions, guardrails, participation monitoring, or teacher-authored prompts, is your school prepared to fund first?
Sources
- Generative AI without guardrails can harm learning: Evidence from high school mathematics
- The Effects of Course-Integrated AI Tutoring on Student Performance and Engagement: A Randomized University Trial
- AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting
- Percentage of Households with Internet Connection Increased to 48.8 percent in 2024 (2024 National ICT Household Survey)
- What the research shows about generative AI in tutoring
Sources — external references open in a new tab.
