All posts
5 min read

Do AI Detectors Actually Work? What the Research Says

AI detectors are treated as evidence in classrooms and newsrooms, but the peer-reviewed data shows accuracy that swings wildly by model, text length, and who wrote it. Here's what the studies actually found.

If you've been flagged by an AI detector, or you're deciding whether to trust one, the honest answer is more complicated than either side of the argument usually admits. Detectors are not random number generators — they measure something real. They are also not the evidence they're often treated as, and the published research is remarkably consistent about why.

We sell a tool in this category, which means you should read what follows with appropriate suspicion. So everything below is sourced to research we didn't run, and we've included the findings that make our own product's category look worse.

What a detector is actually measuring

No detector can see where text came from. There's no watermark in a ChatGPT paragraph and no metadata riding along with it. What a detector does instead is measure statistical properties of the words themselves and estimate how likely a machine was to have produced that particular sequence.

Two properties do most of the work:

Perplexity is a measure of how surprising each word is given the words before it. Language models are built to pick high-probability continuations, so their output tends to have low perplexity — few surprises. Human writing wanders more.

Burstiness measures variation across a passage: sentence lengths, rhythm, how much the complexity rises and falls. Human writing is bursty. We follow a twenty-six-word sentence with a four-word one. Model output tends to cluster in a narrower band.

The important consequence: a detector isn't asking "was this written by AI?" It's asking "is this text unusually predictable?" Those aren't the same question, and the gap between them is where every false positive lives.

Finding 1: The false-positive problem is concentrated, not random

The most-cited study in this space comes from a Stanford group publishing in Patterns. They ran seven GPT detectors against essays written by non-native English speakers for the TOEFL exam — all human-written, verified.

The average false-positive rate was 61.3%. All seven detectors unanimously misclassified 19.8% of the essays as AI-generated. At least one detector flagged 97.8% of them.

Against essays by native English speakers, the same detectors performed close to perfectly.

The mechanism is not mysterious once you know what's being measured. Someone writing in a second language tends to use a more constrained vocabulary and more conventional sentence construction — not because they're less capable, but because that's what writing in an acquired language looks like under exam conditions. That profile is low perplexity. The detector sees predictable text and reports what it was built to report.

This is the single most important thing to understand about detectors: the errors are not distributed evenly. They fall hardest on non-native speakers, and separate research indicates that formal academic register and some neurodivergent writing styles attract elevated flag rates for the same underlying reason.

Finding 2: Vendors are trading sensitivity against fairness

Turnitin, the detector most widely deployed in universities, reports a document-level false-positive rate of 0.51%. Its chief product officer has also stated that the tool catches roughly 85% of AI-generated writing — and that the remaining 15% is allowed to pass deliberately, a threshold set to keep false accusations rare.

That admission is more useful than it first appears. It tells you detection sensitivity is a dial, not a fixed capability. Every vendor is choosing a point on a trade-off curve between catching more AI and wrongly accusing more people. A detector advertising very high catch rates is telling you, implicitly, where it set that dial.

Finding 3: Accuracy degrades as models improve

Detector performance varies enormously depending on which model produced the text. Independent round-ups have found detection sensitivity that is respectable against output from older models and substantially weaker against newer releases.

This is structural. Detectors are trained on the statistical fingerprint of models that existed when the detector was trained. Every new frontier model ships with a slightly different fingerprint, and there's a window before detectors retrain. The window closes — but a new one opens with the next release.

The practical consequence is that any accuracy claim without a date and a named model is close to meaningless. That includes accuracy claims made by humanizer tools.

Finding 4: Short text is not scoreable

Below roughly 150 words, there isn't enough text for the statistical measures to mean much. Perplexity and burstiness are distributional properties; they need a distribution.

Most detectors will still return a confident-looking percentage on a single paragraph. That number is noise wearing a lab coat. If you're evaluating a detector result on a short passage, the length alone is grounds to discount it.

So — do they work?

Sorting the evidence honestly:

What they do reasonably well. Identifying unedited, long-form output from models the detector was trained against, in aggregate, when you care about trends rather than individual verdicts.

What they do badly. Individual determinations of guilt. Anything short. Text from very recent models. And — most seriously — text written by non-native English speakers, where the error rate is high enough that the tool is arguably unfit for that purpose.

What they cannot do at all. Prove authorship. A detector produces a probability estimate about textual style. It's circumstantial at best, and researchers have argued that as models converge on the range of ordinary human variation, reliable detection may not be achievable even in principle.

If you've been flagged

A few things worth knowing, in order:

  1. A score is not evidence. Ask which detector, what score, and on how much text. If it was under 150 words, say so.
  2. Ask about the false-positive rate. If English isn't your first language, the Stanford figures are directly relevant and citable in your own defence.
  3. Show your process. Draft history, version history in your document editor, notes, and outlines are far stronger evidence of authorship than any detector output is of the opposite.
  4. Ask what the policy actually says. Many institutions require corroborating evidence beyond a detector score, precisely because of the research above.

Where we stand

We build a tool that rewrites AI-drafted text to read more naturally. We won't tell you it makes anything undetectable, because that claim isn't true for any tool in this category and the research above is exactly why.

What we'll say is narrower: the patterns detectors measure are real properties of writing, and text that varies its rhythm and avoids template phrasing reads better to humans regardless of what any detector says about it. That's the part that holds up.

Try it on a paragraph — 500 words a day free, no account needed.

Open the editor