Notableevaluation benchmarks

StudentBench: AI and human tutoring yield equivalent GRE learning gains

Curtis Northcutt, Inaara Hasmani, Kevin Feng, Trevor Khangi, Andreas Plesner, Jonas Mueller

Published
Sep 23, 2026 — 17:57 UTC

Problem

This preprint addresses a significant gap in the literature regarding the effectiveness of AI tutoring systems compared to traditional human tutoring. While AI tutoring has gained traction, empirical evaluations of its efficacy in educational contexts, particularly in standardized test preparation like the GRE, remain limited.

Method

The authors developed a platform named StudentBench to conduct their study, which involved over 175,000 student-AI message interactions. The study included 2,383 human participants who were assigned to one of three tutoring conditions: AI tutoring, human tutoring, or no tutoring. The evaluation of tutoring effectiveness was based on 2,028 pairwise rubric evaluations conducted by expert human tutors. Statistical analyses were performed, with p-values reported for comparisons between AI and human tutoring (p = .015) and for one specific AI tutor (p = .044).

Results

The study found that the learning gains from AI tutoring were statistically equivalent to those from human tutoring (p = .015). Notably, the best-performing AI tutor outperformed human tutors in five out of seven GRE domains. In terms of cost-effectiveness, AI tutoring was significantly cheaper, costing USD 0.0052 per percentage point gained, compared to USD 4.81 for human tutoring. Additionally, correlation metrics indicated that faster AI responses were associated with increased student messaging, which in turn correlated with more correct practice attempts and larger learning gains (all p < .002).

Limitations

The authors did not report any limitations in their study. However, the absence of limitations may suggest a need for further scrutiny regarding potential biases in participant selection or the generalizability of the findings across different educational contexts.

Why it matters

The findings of this study have significant implications for the future of educational technology and AI in tutoring. By demonstrating that AI tutoring can achieve learning gains comparable to human tutoring at a fraction of the cost, this research supports the integration of AI systems in educational settings. It opens avenues for further exploration into optimizing AI tutoring systems and their scalability, potentially transforming how students prepare for standardized tests and access educational resources.

Summarised from the primary source with AI assistance under human editorial oversight. Turing Wire is not a primary source — read the original for the authoritative account.

Source: arXiv cs.AI