All posts
6 min read

Why AI Detectors Flag Non-Native English Writers

A Stanford study found AI detectors falsely flagged 61.3% of human-written TOEFL essays, while performing near-perfectly on native speakers. Here's the mechanism behind that gap, and what to do if it happens to you.

In a Stanford study published in Patterns, seven AI detectors were run against essays written by non-native English speakers for the TOEFL exam. Every essay was human-written and verified as such.

The detectors flagged 61.3% of them as AI-generated, on average. All seven agreed on 19.8% of the essays — unanimously wrong. At least one detector flagged 97.8% of the entire set.

Against essays written by native English speakers, the same seven detectors were close to perfect.

That gap is not a bug that got shipped by accident. It follows directly from what these tools measure, and understanding the mechanism is the difference between feeling arbitrarily accused and being able to explain precisely why the tool is wrong.

The mechanism

AI detectors cannot see where text came from. There is no watermark in model output. What they do instead is score how statistically predictable the writing is, using two main measures.

Perplexity captures how surprising each word is given what came before. Language models pick high-probability words by design, so their output has low perplexity.

Burstiness captures variation — sentence length, rhythm, complexity rising and falling across a passage. Human writing tends to be uneven. Model output tends to sit in a narrower band.

Now consider what writing in a second language looks like, particularly under exam pressure. A writer working in an acquired language typically draws on a more constrained active vocabulary, leans on sentence patterns they're confident are correct, and takes fewer syntactic risks. Formal instruction in English as a second language actively teaches conventional structures and standard transitions.

Every one of those adaptations lowers perplexity and reduces burstiness.

The detector sees text that is unusually predictable and reports what it was built to report. It is not detecting AI. It is detecting predictability — and then a human reads that output as an accusation of cheating.

This is a definitional problem, not a tuning problem

It would be convenient if this were a threshold that could be adjusted. It isn't, and that's the uncomfortable part.

The signal that says "this text is predictable" is the same signal whether the predictability comes from a language model optimising for likely next words or from a person writing carefully in a language they learned as an adult. There is no additional feature in the text separating those two cases. A detector tuned to stop flagging second-language writing would have to stop flagging a large share of model output too.

Researchers at TRAILS and elsewhere have made the broader version of this argument: as models improve, their output moves further into the range of ordinary human variation, and the statistical gap detectors rely on narrows. If that holds, reliable detection isn't a problem waiting for a better algorithm.

Who else this affects

Non-native speakers are the most-studied group, but the same mechanism reaches further:

  • Formal academic prose. Discipline-specific conventions produce exactly the standardised phrasing detectors score as predictable. The better you've internalised your field's register, the more you look like a machine.
  • Neurodivergent writers. Research indicates elevated flag rates for some autistic, ADHD, and dyslexic writers, where consistent structural patterns or heavy reliance on learned templates are part of how the writing gets done. This is barely covered anywhere and deserves far more attention than it gets.
  • Anyone taught to write to a rubric. Five-paragraph-essay training optimises for uniform structure. Uniform structure scores as low burstiness.

There's a common thread. The people most likely to be flagged are often those who worked hardest to write "correctly" — and the students with the least institutional standing to contest the accusation.

What Turnitin's own numbers add

Turnitin reports a document-level false-positive rate of 0.51%, well below the Stanford figures. The comparison isn't apples to apples: Turnitin's number is across its general document population, not specifically second-language writing, and independent testing has produced substantially higher rates on targeted samples.

More useful is what Turnitin's chief product officer has said about tuning: the tool catches roughly 85% of AI writing, and about 15% is deliberately allowed through to keep false accusations low.

That's a vendor confirming the trade-off is real and that they chose a point on it. Any detector could catch more AI. It would flag more innocent people to do it.

If you've been flagged

Practical, in order of what actually helps:

1. Get the specifics. Which detector, what score, on how many words. Under roughly 150 words, the score is statistically meaningless regardless of how confident the percentage looks — detectors need a distribution to measure a distribution.

2. Cite the research directly. The Stanford study is peer-reviewed, published in a Cell Press journal, and directly on point if English is not your first language. "The detector was wrong" is weak. "Seven detectors falsely flagged 61.3% of human-written TOEFL essays in a peer-reviewed Stanford study, and I am in exactly that population" is not.

3. Produce your process. Version history in Google Docs or Word, draft files, outlines, notes, browser research history, timestamps. Evidence of the work developing over time is far stronger proof of authorship than a detector score is of the opposite. Going forward, keep drafts — for anyone writing in a second language in an institution that uses detectors, this is unfortunately worth doing routinely.

4. Read the policy. Many institutions require corroborating evidence beyond a detector score. Some prohibit acting on detector output alone. Find out what yours actually says before the conversation.

5. Ask about the appeal route. If the response is that the detector is definitive, that position is not supported by the vendors' own published documentation — including Turnitin's.

For educators reading this

If you use these tools, the research asks something specific of you: a detector score is a prompt to look more closely, not a finding. Applied as a verdict, it produces a disparate-impact problem — the tool is measurably less accurate on international students than on domestic ones, which makes uniform application unfair in effect regardless of intent.

Asking a student to walk you through their draft history takes ten minutes and is both more reliable and more respectful than a percentage.

Where we stand

We build a tool that rewrites AI-drafted text to read more naturally, so we have an obvious interest in you distrusting detectors. Discount accordingly — and note that nothing above is our data. It's linked, published research.

We'll also say the thing that isn't in our commercial interest: if you were falsely flagged and you wrote the text yourself, running it through a humanizer is not the right move. It doesn't establish your authorship, and if the process comes up later, the edit is harder to explain than the original draft. Your draft history is the defence. Use it.

And if you want the fuller picture on detector accuracy generally, we've written that up in more detail.

Try it on a paragraph — 500 words a day free, no account needed.

Open the editor