On June 1, 2026, Stanford Law School published a blind-evaluation study, “Law Professors Prefer AI Over Peer Answers.” Sixteen law professors across U.S. law schools answered 40 contract-law questions, then blindly graded anonymized responses without knowing whether they came from AI systems or fellow professors. Across nearly 3,000 blind comparisons, AI won 75% of head-to-head matchups.
This is not another “AI beats a benchmark” story. The graders were law professors — a population professionally trained to pick apart arguments. The task was tutoring, not exam scoring. And AI answers were flagged as pedagogically harmful or misleading only 3.5% of the time, versus 12% for professor-written answers. The study was led by Stanford Law professor Julian Nyarko, head of the Legal Innovation through Frontier Technology Lab (liftlab), with liftlab researcher Alejandro Salinas as first author and co-authors including Sarath Sanga of Yale Law plus colleagues from NYU, the University of Chicago, and other institutions.
How the Study Worked
Three design choices carry the weight. First, the questions were deliberately mundane: 40 contract-law prompts of the kind students actually ask after class or during office hours — not competition problems. Second, the grading was blind: professors wrote their own answers, then scored anonymized responses with no indication of whether an answer came from AI or a colleague. Third, for fairness, AI responses were calibrated to match the length and structure of human answers, so graders could not guess the source from formatting, and multiple evaluation methods were used as cross-checks.
Sixteen professors and roughly 3,000 anonymized comparisons is not a huge sample, but the blind design controls for brand halo and identity bias — exactly what most “AI beats X” studies are missing. The length-and-structure calibration matters too: without it, graders tend to reward verbosity or penalize formatting they do not recognize, which would tilt the comparison toward whichever side happened to write longer.
The Core Numbers
- AI beat professor-written answers in 75% of head-to-head matchups
- AI performed comparably to the best human instructor in the study
- AI answers were flagged as pedagogically harmful or misleading 3.5% of the time, versus 12% for peer-written answers
The third number gets the least attention and may matter most. When professors graded blind, they had fewer safety concerns about AI answers than about their colleagues’ answers. That runs directly against the popular intuition that AI will mislead students.
Which Systems Were Tested
The study examined commercial tutoring systems and Google’s NotebookLM, with varying performance levels across systems. The press release does not name a full model list; the paper is available as a PDF and via SSRN (abstract 6849678). The notable detail: these are off-the-shelf products, not specially tuned models. Any student can use the same tools today, which makes the finding feel a lot less hypothetical than a lab evaluation.
Limits and the Authors’ Own Caveats
Nyarko was careful to stress that the study evaluated answer quality only — “how to implement these tools to most effectively improve student learning is still an open question.” In plain terms: AI answers well; that does not mean students using AI learn more. The blind design also has boundaries. All 40 questions sit in a single subject, contracts; 16 professors cannot represent all of legal education, let alone other disciplines.
The study landed hard on Hacker News after publication (417 points, 356 comments), and most of the argument there was about the same gap: answer quality versus educational outcome.
What It Means for EdTech
Three observations. First, introductory-level Q&A is likely the first place AI gets adopted at scale in higher education — it is high-frequency, time-consuming for instructors, and has a blind-testable quality signal. Second, once AI answers beat the instructor’s own in blind grading, the institutional question shifts from “should we ban this” to “how do we redesign tutoring and assessment” — a ban will not stop tools every student already has. Third, the appearance of NotebookLM — a system grounded in course materials — in a rigorous study means “AI plus your own materials” now has measurable evidence behind it, not just marketing copy. The study does not settle whether AI belongs in the classroom, and its authors say so plainly. What it does is move the burden of proof: from here on, the case against AI tutoring has to be argued with data, not intuition.
Sources
- AI outperforms law professors in Stanford Law study — Stanford Law School
- Law Professors Prefer AI Over Peer Answers — Stanford Law School Publications
- Salinas et al. paper PDF — Stanford Law School
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
