What an AI Detection Score Actually Means — A Guide for Educators
Vanderbilt, Waterloo, Yale and Curtin all switched off Turnitin's AI detector, and the MLA-CCCC task force advises against treating detector output as proof. Here is what a score can support, what it cannot, and how to handle a flag fairly.
An AI detection score is a probability estimate about text patterns, not a finding about a person. That distinction is the whole of this article, and it is the reason Vanderbilt University disabled Turnitin's AI detector in August 2023 after its own testing flagged entirely human-written text as 100% AI-generated. Yale, the University of Waterloo and Curtin University have since made the same call. The MLA-CCCC Joint Task Force on Writing and AI reached a compatible conclusion from a different direction: detection output should not carry an academic integrity case on its own.
None of that means detectors are worthless. It means the score is a prompt to look closer, and nothing beyond that.
A disclosure before we go further, because it affects how you should read this. We build a tool that rewrites AI-generated text to read more naturally, so we sit on the commercial side of this argument and you should weigh what follows accordingly. Every claim below is sourced to university announcements, peer-reviewed research, and professional-body guidance that we did not produce. We have also included the findings that make our own product category look worse, because a guide for educators that omitted them would be worthless to you.
Key takeaways
- Vanderbilt disabled Turnitin's AI detector in August 2023, citing a lack of transparency and the fact that a 1% false-positive rate applied to its ~75,000 annual submissions implies roughly 750 papers wrongly flagged. [1]
- The University of Waterloo discontinued the feature as of September 2025, announced by the Associate Vice-President, Academic after consultation across faculties. [2]
- Curtin University disabled AI detection from 1 January 2026, keeping ordinary text-matching in place — only the Gen-AI detection feature was switched off. [3]
- Yale's Poorvu Center states plainly that "the AI detection feature of Turnitin is currently disabled." [4]
- A **2023 Stanford study in Patterns found seven detectors misclassified human-written TOEFL essays at an average 61.3% false-positive rate**, versus near-perfect accuracy on native-English essays. [5]
- The MLA-CCCC Joint Task Force cautions that false accusations may disproportionately affect marginalized groups and urges support-based rather than punitive approaches. [6]
- A peer-reviewed June 2026 study of four detectors concluded an AI score should serve only a signaling function to prompt a closer look. [7]
What this guide covers
- What a detector actually measures
- Why the errors are not evenly distributed
- What the universities that switched it off concluded
- What a score can and cannot support
- A fair process when something is flagged
- Assessment design that does not depend on detection
- Honest limitations, then a FAQ and full sources
What a detector actually measures
No detector can see the origin of a piece of text. There is no watermark in a ChatGPT paragraph and no metadata traveling with it. What these tools measure is how statistically predictable the writing is.
Two properties do most of the work. Perplexity captures how surprising each word is given the words before it; language models select high-probability continuations, so their output tends to run low. Burstiness captures variation across a passage — sentence length, rhythm, the rise and fall of complexity. Human writing tends to be bursty. Model output clusters in a narrower band. (We have written a fuller explainer on what burstiness actually measures if the term is new to you.)
So the question a detector answers is not "did a machine write this?" It is "is this text unusually predictable?" Those are different questions, and the space between them is where every false positive lives.
This also explains a pattern you may have noticed in your own marking. Highly structured, formal, carefully edited academic prose is supposed to be predictable. Clear topic sentences, conventional transitions, controlled vocabulary — the qualities we teach and reward are the qualities that lower perplexity.
Why the errors are not evenly distributed
This is the part that matters most for fairness, and it is well documented.
A 2023 Stanford study published in Patterns ran seven GPT detectors against TOEFL essays written by non-native English speakers, all verified human-written. The average false-positive rate was 61.3%. All seven detectors unanimously misclassified 19.8% of the essays, and at least one flagged 97.8% of them. Against essays written by native English speakers, the same detectors performed close to perfectly. [5]
The mechanism is not mysterious. Writing in an acquired language, particularly under time pressure, tends to produce a more constrained vocabulary and more conventional sentence construction. That is a low-perplexity profile. The detector reports what it was built to report. We have traced why this happens to non-native English writers specifically in more detail, including why it is a structural property of text-only detection rather than a tuning problem.
The same logic extends to other groups: students writing in a formal register, students in technical and legal disciplines, students who edit heavily, and some autistic and neurodivergent writers. In Matter of Newby v. Adelphi University — decided in New York Supreme Court on 26 February 2026 — a paper by a freshman with Level 2 Autism Spectrum Disorder returned a 100% Turnitin AI score while two other detectors read the same text as human-written. The court annulled the violation and ordered the record expunged. [8]
The practical consequence for you: a detector flag is not a random error you can treat as rare background noise. It is a systematically biased error that lands hardest on the students least equipped to contest it.
What the universities that switched it off concluded
The institutional record is more useful than any vendor's accuracy claim, because these are organizations that ran the tool on their own students and then published their reasoning.
| Institution | Decision | Effective | Stated reasoning |
|---|---|---|---|
| Vanderbilt University | Disabled Turnitin AI detection | August 2023 | No transparency into how determinations are made; ~750 of 75,000 annual papers implied wrongly flagged at a 1% FPR; internal tests returned 100% AI on human text; documented bias against non-native speakers [1] |
| University of Waterloo | Discontinued the feature | September 2025 | Decision taken after consultation across Undergraduate Associate Deans, Academic Deans and the EdTech Steering Committee [2] |
| Curtin University | Disabled Gen-AI detection; text-matching retained | 1 January 2026 | Fostering "trust and clarity within a modern academic culture"; assessments to be "secure, fair, relevant and future-ready" [3] |
| Yale University | AI detection feature disabled | Current as of 2026 | "The AI detection feature of Turnitin is currently disabled" [4] |
Two cautions about how this trend gets reported. First, you will see claims that "more than 50 universities" have disabled AI detection, often with Johns Hopkins, Northwestern, NYU and others named. We could not verify most of those against official institutional sources, so we are not repeating them as fact — the four above are the ones with primary documentation we could confirm, and we maintain a verified tracker of which universities disabled AI detection that separates the confirmed cases from the merely repeated ones. Second, disabling a detector is not the same as permitting undisclosed AI use. Curtin kept text-matching running and its academic integrity procedures unchanged. [3]
What a score can and cannot support
The June 2026 study in the International Journal for Educational Integrity tested four detectors against 160 academic papers of over 4,000 words, evenly split between ESL human-written, AI-generated, hybrid, and humanised AI text. Only one of the four produced results the researchers considered satisfactory. Their conclusion is the operative sentence for practitioners: an AI score can serve a signaling function to prompt a closer look, and no more. [7]
That gives a workable division:
A score can reasonably support:
- A decision to read the submission more carefully yourself
- A conversation with the student about their process and their argument
- A pattern observation across a body of work over a term
- A prompt to review whether the assessment design is doing its job
A score cannot reasonably support:
- A misconduct finding on its own
- A grade penalty applied before the student has responded
- An accusation communicated as settled fact
- A conclusion about a student whose first language is not English, without accounting for the documented bias
Northwestern's academic integrity guidance reportedly states that detection results alone are insufficient for a misconduct finding, and Johns Hopkins is reported to treat detection as advisory only — a hit can start a conversation but not a charge. We flag both as secondary-sourced rather than confirmed from institutional pages. [9]
A fair process when something is flagged
The Adelphi ruling is instructive precisely because the court did not need to decide whether Turnitin works. It found the process defective on grounds any institution can avoid. The violation was annulled because contrary evidence the student submitted was never meaningfully addressed, the university failed to comply with its own Code of Conduct, the student was denied an advisor, and the same administrator who issued the finding also decided the appeal. [8]
A defensible sequence:
- Read the work yourself before doing anything else. Form your own view of whether it matches the student's demonstrated ability and the assignment.
- Open with a question, not a charge. "I would like to talk through how you approached this" preserves the relationship and your options. An accusation you later withdraw does lasting damage.
- Ask for process evidence. Version history in Google Docs or Word, dated drafts, notes and outlines. This is far stronger evidence in both directions than any score. It is also what we tell students to assemble in our guide to answering a false AI accusation — worth reading if you want to know what a well-prepared student will bring to the meeting.
- Discuss the substance. A student who can defend the thesis, explain why a source was chosen, and say what they cut is telling you something a detector cannot.
- Record the detector, the version, and the score — and note its known error profile, particularly if the student is a multilingual writer.
- Address contrary evidence explicitly if the student produces it. Ignoring it was decisive at Adelphi. [8]
- Separate the roles. Whoever makes the initial finding should not decide the appeal.
- Permit an advisor where your policy allows it.
Assessment design that does not depend on detection
The MLA-CCCC Joint Task Force urges approaches that support students rather than punish them, and points to assignment design as the more durable answer. [6] Detection is a race you do not control; assessment design is one you do.
Approaches colleagues report as effective:
- Assess the process, not only the artifact — outlines, annotated bibliographies, draft submissions and reflections, each carrying marks.
- Anchor tasks in the specific — this seminar's discussion, a local dataset, a text only your cohort read. Generic prompts get generic answers from any source.
- Build in an oral component, even briefly. Two minutes explaining a choice reveals more than a score.
- State your AI policy in writing, per assignment. Harvard, Oxford and the University of Michigan have reportedly moved to explicit disclosure language rather than blanket bans. Ambiguity is where most disputes begin, and "I did not know that counted" is often true.
- Distinguish grammar support from generation. A multilingual student using a grammar checker is not doing what a student generating an essay is doing, and a policy that fails to separate them will be applied unevenly.
Honest limitations
We are a writing-tools company, not a teaching-and-learning centre, and this is general information rather than institutional or legal advice. Your own policies govern.
Some other things worth stating plainly:
- The "50+ universities" figure is not something we could verify. Four institutions above have primary documentation. Treat larger counts circulating in this space, including on competitor sites, as unconfirmed.
- Newby is a single New York trial-level decision. It is persuasive and widely reported, not binding nationally.
- Detector accuracy moves quickly. The 2023 Stanford figures describe the detectors of that period. Turnitin has since revised its model, and the 2026 study found wide variance between tools — so which tool produced a score is a material fact, and any number should be cited with its date. We keep the published accuracy data we rely on, with dates and sources, on our detector test results page, and a broader survey of the evidence in do AI detectors actually work?
- Yale's page states the feature is disabled but gives no reasoning or date, so we have not attributed motives to that decision.
- We have no first-party testing of our own to offer here. Where we cite accuracy figures, they come from researchers and institutions, not from us.
FAQ
Can I use a detector score as evidence of misconduct? Not on its own. The June 2026 study frames detector output as a signal to look closer rather than proof [7], the MLA-CCCC task force cautions against relying on it [6], and the Adelphi court annulled a violation where contrary evidence went unaddressed [8].
Is Turnitin's AI detection accurate? Turnitin has reported a document-level false-positive rate of 0.51% and has said the tool is tuned to let roughly 15% of AI writing pass in order to keep false accusations rare. Those are vendor figures. Vanderbilt's objection was partly that even a 1% rate implies around 750 wrongly flagged papers a year at its volume. [1]
Why do multilingual students get flagged more often? Writing in an acquired language tends to produce constrained vocabulary and conventional constructions, which reads as low perplexity — the pattern detectors are built to flag. The 2023 Stanford study measured a 61.3% average false-positive rate on human-written TOEFL essays. [5]
Should my institution disable AI detection entirely? That is a decision for your teaching-and-learning and legal teams, and reasonable institutions have gone both ways. Curtin's approach is worth noting as a middle path: Gen-AI detection off, text-matching and integrity procedures retained. [3]
A student says a detector cleared their work. Does that prove anything? Only weakly. Detectors disagree constantly — the 2026 study found only one of four tools performed satisfactorily [7]. It is not proof of innocence, but it is contrary evidence, and Adelphi established that failing to address contrary evidence can render a finding unreasonable. [8]
What is the single most useful thing to ask for? Version history. A timestamped record of a document being built incrementally is stronger evidence than any detector output, in either direction.
Does disabling detection mean permitting undisclosed AI use? No. Curtin retained text-matching and left its academic integrity procedures in place. [3] The two decisions are separate.
Sources
- "Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector," Vanderbilt University Brightspace, 16 August 2023 — the institution's own published reasoning, cited for the transparency objection, the 750-papers calculation, and the internal tests that returned 100% AI on human text. A first-hand account from a university that ran the tool on its own students.
- "Waterloo discontinuing the use of AI detection tool Turnitin.com," Office of the Associate Vice-President, Academic, University of Waterloo — official announcement, cited for the September 2025 effective date and the consultation process. Used because it is the primary record rather than press coverage of it.
- "Update on Turnitin AI-Detection Tool," Curtin University — official notice, cited for the 1 January 2026 effective date, the stated reasoning, and the important detail that text-matching was retained.
- Poorvu Center for Teaching and Learning, Yale University — Turnitin instructional-tools page — cited only for the verbatim statement that the AI detection feature is disabled. Included precisely because it is Yale's own page; no reasoning is attributed since the page gives none.
- **Liang et al., "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023** — peer-reviewed primary source for the 61.3% average false-positive rate and the 97.8% figure. Cited as the study itself rather than a summary.
- MLA-CCCC Joint Task Force on Writing and AI, "Generative AI and Policy Development" and "Academic Integrity and Assignment Design" — guidance from the relevant professional bodies for writing instruction, cited for the disproportionate-impact caution and the assignment-design recommendation.
- **"Who wrote this? Evaluating the reliability of AI detection tools in higher education," International Journal for Educational Integrity (Springer), June 2026** — Vrije Universiteit Brussel. Peer-reviewed and current; the source for the four-detector comparison, the 160-paper methodology, and the "signaling function" conclusion.
- Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021 (N.Y. Sup. Ct., 26 February 2026) — the court record, cited for the procedural findings and the expungement order. Corroborated in reporting by CBS News New York and analyses by Liebert Cassidy Whitmore and LLF National Law Firm.
- Secondary aggregator coverage (AI Weekly and similar) — the source for the Johns Hopkins and Northwestern characterisations and the "50+ universities" figure. Flagged explicitly as unverified against institutional sources, and identified here so readers can weigh it accordingly rather than mistake it for confirmed fact.
Try it on a paragraph — 500 words a day free, no account needed.
Open the editor