Why AI Detector Bias Against ESL Writers Cannot Be Fixed
A March 2026 arXiv paper proves the bias against multilingual writers is a mathematical property of text-only detection, not an engineering flaw. Better detectors will not solve it. Here is the proof, the evidence, and what it means for students and institutions.
The bias in AI detectors against multilingual writers is not a bug that a better model will eventually fix. A March 2026 paper by Nathan Garland frames detection as a composite hypothesis problem and shows that any text-only, one-shot detector with meaningful detection power will necessarily produce false accusations among writers whose prose overlaps statistically with model output. The constraint, in the paper's words, "cannot be overcome by better detector engineering or technology." [1]
That changes the argument. For three years the standard defence of these tools has been that accuracy is improving and the equity problem will shrink with it. If the mathematics holds, it will not shrink. It is structural. This page assumes you already know why detectors flag non-native English writers at all; what follows is the proof that the problem has no engineering exit.
Our disclosure first, because it matters here more than usual: we sell a tool that rewrites AI-generated text, so a reader who distrusts detectors is a reader who might buy from us. That is a real conflict and you should weigh this page accordingly. It is also why the argument below rests on a peer-reviewed study, an arXiv preprint you can read in full, a court record, and universities' own published decisions — not on our assertions. We have included the part that cuts against our own product, too: if you are a multilingual writer facing an accusation over work you actually wrote, a humanizer is the wrong tool and will make your position worse.
Key takeaways
- Garland (arXiv:2603.20254, 11 March 2026) proves via a subgroup mixture bound that disparate impact is "logically independent of AI model quality" — better engineering cannot remove it. [1]
- The **2023 Stanford study in Patterns measured a 61.3% average false-positive rate across seven detectors on human-written TOEFL essays; at least one detector flagged 97.8%** of them. Native-English essays were classified near-perfectly. [2]
- The mechanism is perplexity: writing in an acquired language produces more predictable word choices, which is precisely the signal detectors are built to flag. [2]
- Vanderbilt cited bias against non-native speakers as an equity concern when it disabled Turnitin's AI detector in August 2023. [3]
- A June 2026 peer-reviewed study of four detectors concluded a score should only serve a signaling function to prompt a closer look. [4]
- In Matter of Newby v. Adelphi University (26 February 2026), a New York court annulled a violation built on a 100% Turnitin AI score and ordered the record expunged. [5]
- Garland's practical recommendation matches the courts and the professional bodies: detection scores should not be sole evidence in misconduct proceedings. [1]
What this page covers
- Why predictable writing is not machine writing
- The structural argument, in plain language
- What the empirical evidence shows
- Who else this catches
- For multilingual students: what actually helps
- For institutions: what follows from a structural limit
- Honest limitations, then a FAQ and full sources
Why predictable writing is not machine writing
A detector cannot see where text came from. There is no watermark in a ChatGPT paragraph. What these tools measure is perplexity — how surprising each word is given the words before it — and burstiness, the variation in sentence length and complexity across a passage.
Language models pick high-probability continuations, so their output tends to score low on both. The inference is that low-perplexity text is probably machine-written. If either term is unfamiliar, our explainer on burstiness in AI writing covers what these measurements actually capture, and do AI detectors actually work? surveys how well the inference holds up in practice.
Here is the failure. Writing in an acquired language also produces low-perplexity text, for reasons that have nothing to do with machines. A writer working in a second language, particularly under exam conditions, tends to reach for the constructions they are most confident in: a more constrained vocabulary, more conventional sentence patterns, safer transitions. That is not a deficiency. It is what competent writing in an acquired language looks like when the stakes are high.
The detector sees predictability and reports predictability. It has no way to distinguish "predictable because a model generated it" from "predictable because the writer was being careful in their third language."
The structural argument, in plain language
Garland's contribution is to show that this is not a temporary limitation. [1]
The setup: a detector has to separate two populations using a single number derived from the text. Human writing is not one distribution — it is a mixture of many, one per subgroup, and those subgroups differ in how predictable their prose tends to be. Some human subgroups sit statistically closer to model output than others.
The subgroup mixture bound connects a detector's error rates to that demographic variation. Once some human subgroup overlaps the machine distribution, any threshold you choose forces a trade-off: set it loose enough to catch AI text reliably, and you necessarily sweep in members of the overlapping human subgroup. Tighten it to protect them, and detection power drops.
Two consequences follow, and they are the reason this paper matters more than another accuracy benchmark:
First, it is independent of model quality. The constraint "aris[es] from population diversity" and is "logically independent of AI model quality." [1] A better-trained detector does not dissolve the overlap, because the overlap is a fact about human writing, not about the detector.
Second, the assessor does not know the individual's distribution. [1] To judge whether this student's low-perplexity paragraph is unusual, you would need to know their personal writing baseline. A one-shot detector, handed a single document with no history, does not have that. It compares the text to a population average that may not describe the writer at all.
That second point is the one worth carrying into any appeal or policy meeting. It explains why version history and prior coursework are stronger evidence than any score: they supply exactly the individual baseline the detector structurally lacks.
What the empirical evidence shows
The theory predicts a pattern the data had already found.
The 2023 Stanford study published in Patterns ran seven GPT detectors against TOEFL essays written by non-native English speakers, all verified human-written. The average false-positive rate was 61.3%. All seven detectors unanimously misclassified 19.8% of the essays. At least one flagged 97.8% of them. Against essays by native English speakers, the same detectors performed close to perfectly. [2]
Read those two results together. The tools were not broadly inaccurate. They were accurate for one population and close to useless for another — which is the signature of a structural bias rather than general noise.
The institutional record points the same way. When Vanderbilt disabled Turnitin's AI detector in August 2023, documented bias against non-native English speakers was among the equity concerns its Instructional Technologies team raised, alongside internal tests that returned 100% AI on human-written text. [3] Vanderbilt is one of four institutions in our verified tracker of universities that disabled AI detection, each confirmed against the university's own published page. The June 2026 study in the International Journal for Educational Integrity tested four detectors on 160 papers — a quarter of them ESL human-written — and found only one performed satisfactorily, concluding that a score should prompt a closer look and nothing more. [4]
| Finding | Source | Date | What it establishes |
|---|---|---|---|
| 61.3% average FPR on human TOEFL essays; 97.8% flagged by at least one detector | Stanford, Patterns [2] | 2023 | The disparity is large and measurable |
| Disparate impact is mathematically structural; independent of model quality | Garland, arXiv:2603.20254 [1] | Mar 2026 | Better detectors will not fix it |
| Only 1 of 4 detectors satisfactory; score is a signal, not proof | IJEI (Springer) [4] | Jun 2026 | Tool choice is a material fact |
| Bias cited as equity concern when disabling the tool | Vanderbilt [3] | Aug 2023 | Institutions acted on this evidence |
| Violation annulled where contrary evidence went unaddressed | Newby v. Adelphi [5] | Feb 2026 | Courts scrutinise process |
Who else this catches
The mixture bound is about statistical overlap, not about nationality, so multilingual writers are not the only affected group. Any subgroup whose writing runs more predictable than average sits in the same position:
- Students writing in a formal academic register — the clarity and conventional structure we teach are low-perplexity features.
- Technical, legal and scientific writers, where standardised phrasing is a professional requirement.
- Heavy editors. Polishing prose tends to regularise it. Recent coverage of this research has noted the paradox directly: careful revision moves human text toward the profile detectors flag.
- Some autistic and neurodivergent writers. In the Adelphi case, the flagged student had a Level 2 Autism Spectrum Disorder diagnosis; his paper returned 100% AI on Turnitin while two other detectors read it as human-written. The court annulled the finding and ordered the record expunged. [5]
For multilingual students: what actually helps
Direct advice, including the part that is against our commercial interest.
Do not run disputed work through a humanizer — ours included. If you wrote the paper, rewriting it now destroys the match between your submitted text and your version history, which is the single strongest evidence you have. It also looks like tampering with work under dispute. There is no version of this that improves your position.
Preserve process evidence, starting now and not only when accused. Draft in a tool that keeps version history — Google Docs version history, or Word with Track Changes and OneDrive versioning. This directly supplies the individual baseline Garland shows a detector structurally lacks. [1]
Keep the surrounding material. Outlines, notes, annotated sources, earlier graded work in the same course. Prior coursework is unusually valuable for multilingual writers because it establishes your genuine register over time.
Ask which tool flagged you. Detector quality varies enormously — only one of four tested performed satisfactorily in the June 2026 study. [4] Which tool produced the score is a material fact, not a detail.
Lead with your evidence, then the research. Present your process evidence first. Bring the 61.3% figure [2] and Garland's structural argument [1] as supporting context. The research explains why a flag is unreliable; it does not by itself prove that you did not do it, and panels can tell the difference.
Do not overstate. If you used a grammar checker, say exactly that and name it. Precision protects you; vague admissions do not.
For institutions: what follows from a structural limit
If the bias cannot be engineered away, then "wait for better tools" is not a policy. Some implications, expanded in our companion guide on what an AI detection score actually means for educators:
A threshold is an equity decision, not a technical one. Choosing a sensitivity setting decides which population absorbs the errors. That belongs with the people accountable for equity outcomes, not solely with a vendor default.
Supply the missing baseline. Since the structural gap is the absence of individual writing history [1], assessment that generates that history — drafts, outlines, reflections, in-class writing — fixes the actual problem rather than the symptom.
Convert rates into headcounts. Vanderbilt's most persuasive move was arithmetic: a 1% false-positive rate against roughly 75,000 annual submissions implies around 750 students wrongly flagged in a year. [3] Then ask which students those are likely to be.
Align with where the evidence already points. Garland recommends detection scores not serve as sole evidence [1]; the June 2026 study calls a score a signaling function [4]; the MLA-CCCC Joint Task Force cautions that false accusations may disproportionately affect marginalized groups; and the Adelphi court annulled a finding where contrary evidence went unaddressed and one administrator both issued the finding and denied the appeal. [5] These converge.
Honest limitations
- Garland's paper is an arXiv preprint, submitted 11 March 2026. It is a formal argument you can read and check, but we could not confirm it has completed peer review — unlike the Patterns and IJEI studies, which have. Weigh it as a rigorous argument rather than a settled result.
- A structural limit is not a claim that detectors are useless. The mathematics says errors are unavoidable and unevenly distributed, not that the signal is meaningless. Pangram, for instance, reported markedly better results than its competitors in the June 2026 comparison. Better tools exist; the paper's claim is that no tool escapes the trade-off entirely.
- The 2023 Stanford figures describe the detectors of that period. Vendors have revised their models since, and the 2026 study found wide variance between tools. Cite the study with its date rather than treating 61.3% as a current universal rate.
- We could not independently verify some widely circulated claims, including the "more than 50 universities" figure and reported ESL suspension-risk differentials. They are omitted here rather than repeated because they happen to favor our argument.
- This is general information, not legal or institutional advice. If you face suspension, expulsion, or visa consequences, consult a student defence attorney. The stakes for international students are materially higher, and this page is not a substitute for representation.
- We have no first-party testing to offer on this question. Every figure above comes from researchers and institutions, not from us.
FAQ
Will better AI detectors eventually fix the bias against ESL writers? Garland's paper argues no: the constraint arises from population diversity and is "logically independent of AI model quality," so it cannot be overcome by better engineering. [1]
Why do AI detectors flag non-native English writers so often? They measure predictability. Writing in an acquired language tends to use constrained vocabulary and conventional constructions, which reads as low perplexity — the same signal model output produces. The 2023 Stanford study measured a 61.3% average false-positive rate on human-written TOEFL essays. [2]
Does a low perplexity score mean my writing is bad? No. It often means the opposite. Clear, conventional, carefully edited prose is predictable by design, which is why polished human writing and formal academic register also attract flags.
I am an international student and I have been flagged. What should I do first? Preserve your version history before editing anything, ask which tool flagged you and which passages, and gather prior coursework establishing your baseline. Then request the appeal process in writing. Our full guide to proving you wrote it yourself walks through the first 24 hours and how to assemble the evidence packet.
Should I use a humanizer to make my writing look less AI-like? Not on disputed work, and we make one of these tools. It breaks the link between your text and your version history and reads as tampering.
Are detectors biased against anyone else? Yes — anyone whose writing runs more predictable than average: formal academic registers, technical and legal writing, heavily edited prose, and some autistic and neurodivergent writers. [5]
Can I use this research in an appeal? As supporting context, after your process evidence. Cite the studies directly [1][2][4] rather than this page, and note that Garland is a preprint.
Sources
- Nathan Garland, "AI Detectors Fail Diverse Student Populations: A Mathematical Framing of Structural Detection Limits," arXiv:2603.20254, submitted 11 March 2026 — the central source for the structural argument, the subgroup mixture bound, the independence from model quality, and the recommendation against sole-evidence use. Cited as the paper itself, and flagged as a preprint rather than peer-reviewed.
- **Liang et al., "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023** — the peer-reviewed empirical foundation for the 61.3% average false-positive rate and the 97.8% figure. Used because it is the primary study, and because its native/non-native contrast is what makes the disparity visible.
- "Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector," Vanderbilt University Brightspace, 16 August 2023 — an institution's own published reasoning, cited for the equity concern about non-native speakers, the internal 100%-AI test results, and the 750-papers arithmetic.
- **"Who wrote this? Evaluating the reliability of AI detection tools in higher education," International Journal for Educational Integrity (Springer), June 2026** — Vrije Universiteit Brussel. Peer-reviewed and current; the source for the four-detector comparison, the ESL-inclusive 160-paper dataset, and the "signaling function" conclusion. Included partly because its finding that one detector performed well keeps this page honest about tool variance.
- Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021 (N.Y. Sup. Ct., 26 February 2026) — the court record, cited for the neurodivergent-writer dimension, the conflicting detector results, and the procedural findings. Corroborated by CBS News New York and analyses from Liebert Cassidy Whitmore and LLF National Law Firm.
Try it on a paragraph — 500 words a day free, no account needed.
Open the editor