All posts
15 min read

Falsely Accused of Using AI? How to Prove You Wrote It Yourself

A New York court threw out a 100% Turnitin AI score in February 2026. Peer-reviewed research puts some detectors' false-positive rates on ESL writing above 60%. Here's what actually clears your name — and what wastes your time.

If you've been accused of using AI on work you wrote yourself, the thing that clears you is almost never a counter-detector score. It's process evidence — version history, dated drafts, notes, the trail showing the document was built rather than pasted. In February 2026 a New York judge annulled an academic integrity violation built on a 100% Turnitin AI score and ordered the student's record expunged. The detector reading wasn't the problem the court fixed. The absence of a fair process was.

That distinction matters, because most advice on this topic points you at the wrong target.

We should say plainly that we sell a text-rewriting tool in this category, so we have an obvious commercial interest in how you think about AI detectors. Everything below is sourced to research and court records we didn't produce, and the practical advice points you away from our product, not toward it. If you're facing an accusation over work you actually wrote, a humanizer is the wrong tool and using one can make your position worse. More on that below.

Key takeaways

  • A New York Supreme Court annulled an AI-based integrity violation in Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021, decided February 26, 2026 — calling the finding "without valid basis and devoid of reason." [1][2]
  • The Adelphi student's paper scored 100% AI on Turnitin, while two other detectors classified the same text as human-written. [2][3]
  • A **2023 Stanford study in Patterns found seven GPT detectors misclassified TOEFL essays by non-native English speakers at an average 61.3% false-positive rate; at least one detector flagged 97.8%** of them. [4]
  • A peer-reviewed June 2026 study from Vrije Universiteit Brussel tested four detectors on 160 papers and found only one produced satisfactory results — meaning detector quality varies enormously, and which tool flagged you is a material fact. [5]
  • Draft/version history is the strongest single piece of evidence you can produce. Counter-detector scores are supporting material at best. [6][7]
  • At a public university, due process entitles you to notice, the evidence, and a chance to respond; FERPA gives you a right to see your own records. [6]

What this guide covers

  1. What to do in the first 24 hours
  2. Why detectors flag human writing
  3. What the courts have actually said
  4. The evidence that works, ranked
  5. What doesn't work (and what backfires)
  6. If you did use AI assistance
  7. For educators and administrators
  8. Honest limitations, then a FAQ and full sources

What to do in the first 24 hours

Move on evidence before you move on argument. Files get overwritten, cloud version histories age out, and memory of your own process fades faster than you'd expect.

  1. Stop editing the document. Don't open it and start "fixing" anything. Editing now muddies the version history that is about to become your best evidence.
  2. Export your version history immediately. In Google Docs: File → Version history → See version history. In Word: check Track Changes and AutoRecover versions, plus OneDrive's version list. Screenshot it and export it.
  3. Ask which tool flagged you, and which passages. You cannot defend text when you don't know which text is in question. Put the request in writing. [7]
  4. Collect the surrounding material — outlines, handwritten notes, highlighted PDFs, library records, browser history from your research sessions, messages where you talked about the assignment.
  5. Ask for your institution's written policy on AI detection evidence and the appeals process, and ask whether you may bring an advisor.
  6. Don't confess to something you didn't do to make it end faster. Admitting to "maybe some AI" when you used none is the most common self-inflicted wound in these cases.

Why detectors flag human writing

No detector can see where text came from. There's no watermark in a ChatGPT paragraph. What a detector measures is how statistically predictable your word choices are — primarily perplexity (how surprising each word is, given the ones before it) and burstiness (how much sentence length and complexity vary across a passage).

Language models pick high-probability continuations, so their output tends to run low on both. The catch is that plenty of human writing does too.

This is why the errors aren't spread evenly. A 2023 Stanford study published in Patterns ran seven GPT detectors against TOEFL essays written by non-native English speakers — all verified human-written. The average false-positive rate was 61.3%. All seven unanimously misclassified 19.8% of the essays. At least one detector flagged 97.8% of them. Against essays by native English speakers, the same detectors performed nearly perfectly. [4]

Writing in an acquired language under exam conditions produces a constrained vocabulary and conventional sentence construction. That profile reads as low perplexity. The detector reports what it was built to report. If this is your situation, why AI detectors flag non-native English writers sets out the mechanism behind that 61.3% figure in full — useful context to have ready if you are asked to explain it.

The same mechanism catches other groups: formal academic register, technical and legal writing, heavily-edited prose, and — as the Adelphi case involved — some autistic and neurodivergent writing styles. [2]

What the courts have actually said

The most useful development for anyone in this position is Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021, decided in New York Supreme Court on February 26, 2026. [1]

Orion Newby, a freshman diagnosed with Level 2 Autism Spectrum Disorder and enrolled in Adelphi's Bridges Program, was accused after a history paper returned a 100% AI score on Turnitin. He submitted contrary detection evidence — two other detectors read the paper as human-written. [2][3]

The court granted his petition, annulled both the violation and the appeal denial, ordered the record expunged, and rescinded the sanction. Its reasoning is what you should take from this: [1]

  • The finding was "without valid basis and devoid of reason," particularly because the student's contrary detection evidence was never meaningfully addressed.
  • The university failed to substantially comply with its own Code of Conduct, including the Student Bill of Rights.
  • Newby was denied an advisor, and the same administrator who issued the finding also decided the appeal — so there was no meaningful review.

Read that list again. Two of the three winning points are about procedure, not about detector accuracy. The court didn't need to rule that Turnitin is unreliable. It ruled that the university couldn't ignore contrary evidence and couldn't have one person serve as both prosecutor and appellate judge.

That's the shape of a successful defense: force the institution to follow its own rules, and put contrary evidence on the record so that ignoring it becomes the error.

The evidence that works, ranked

EvidenceStrengthWhy it carries weightGet it from
Cloud version historyStrongestTimestamped, incremental, hard to fabricate. Shows a document being built over hours or days. [6][7]Google Docs version history; Word/OneDrive version list
Dated draft filesStrongIndependent timeline corroborating the version historyYour drive, email attachments to yourself, LMS submissions
Research artifactsStrongShows the thinking that preceded the text — notes, outlines, annotated sourcesNotebooks, PDF annotations, library/database logs
Your own explanation of the argumentModerate–strongCan you defend the thesis and sources in conversation? Panels weigh this heavilyA meeting, prepared
Prior writing samplesModerateEstablishes your baseline voice and registerEarlier graded work in the same course
Counter-detector scoresWeak on their ownDetectors disagree constantly; useful only to show the flag isn't reproducible [2][5]Run the same text through 2–3 tools, dated screenshots
Institutional/published error ratesSupportingContextualizes the flag; doesn't prove your case by itself [4][5]The studies cited here

The pattern in cleared cases is consistent: students who prevailed did so on writing portfolios, draft histories, and contemporaneous notes. [3][6] Counter-detector evidence mattered in Adelphi not because it proved innocence, but because the university's failure to address it made the finding unreasonable. [1]

Assemble it as one narrative

Don't hand over a pile of files. Build a single packet that tells the story in order: researched → outlined → drafted → revised → submitted, with a timestamp attached to each stage. Add a one-page cover summary. A panel that can follow the timeline in five minutes is a panel that can rule for you.

What doesn't work (and what backfires)

Running your own text through a humanizer. This is the single worst move available to you, and we sell one, so take the warning seriously. If you wrote the paper yourself, rewriting it now destroys the correspondence between your submitted text and your version history — the exact evidence that clears you. It also looks, on inspection, like tampering with disputed work. There is no version of this that improves your position. Where these tools are and are not legitimate is a separate question, and we have answered it directly in is using an AI humanizer cheating?

Leading with "detectors are unreliable." True, well-documented, and unpersuasive on its own. It's an argument about the tool, not about you, and it's the argument every actually-guilty student also makes. Use published error rates as supporting context after your process evidence has done the work — do AI detectors actually work? collects the research if you need to cite it, and our detector test results page keeps the figures dated and sourced.

Relying on one counter-detector score. Detectors disagree with each other routinely; the June 2026 VUB study found that of four tools tested on 160 papers, only one produced results the researchers considered satisfactory. [5] A single clean score from a tool the panel has never heard of proves little. Two or three, dated and screenshotted, at least demonstrate the flag isn't reproducible.

Reconstructing drafts after the accusation. Backdating or recreating a "draft" is misconduct in its own right, and file metadata usually gives it away. If you don't have process evidence, say so honestly and lean on the other categories.

Going in alone. Ask whether you may bring an advisor. In the Adelphi case, being denied one was part of what made the process defective. [1]

If you did use AI assistance

Grammar checking, outlining help, a model that suggested a transition — these fall into a genuinely grey zone, and policies vary enormously by institution and instructor.

Two things are true at once. Overstating your case ("I never touched any AI tool") collapses badly if a chat log or an add-on history surfaces later. And volunteering vague admissions ("I guess I used some AI") when the tool only checked your spelling hands the panel a confession it didn't earn.

The workable path: read your institution's actual written policy first, then describe exactly what you used and for what — "Grammarly for grammar on the final pass," "a model to help outline section three, all prose written by me" — and map that against the policy's own language. Precision protects you. Vagueness doesn't.

If the policy genuinely permitted what you did, that's a policy argument you can win. If it didn't, an honest early account tends to be treated far better than one extracted later.

For educators and administrators

If you're on the other side of this, the research points somewhere specific.

The June 2026 Vrije Universiteit Brussel study tested Pangram, GPTZero, Turnitin, and Copyleaks against 160 academic papers of 4,000+ words, evenly split between ESL human-written, AI-generated, hybrid, and humanized AI text. The researchers concluded that of the four, only one produced satisfactory results. Applied afterward to 1,163 real master's theses, that tool flagged 45.5% for likely AI content — mostly partial use, not wholesale generation. [5]

The authors' own conclusion is the operative one: an AI score can serve a signaling function to prompt a closer look, and nothing more. [5]

Practically, that means: know your tool's published error rate and how it performs on ESL writing; never let one score be the whole case; ensure the person who decides the appeal isn't the person who made the finding; and address contrary evidence a student submits on the record. The Adelphi ruling turned on the last two. [1]

Honest limitations

This is general information, not legal advice, and we're a writing-tools company rather than a law firm. If your case involves suspension, expulsion, visa status, or a professional license, talk to a student defense attorney.

A few other things worth being straight about:

  • Newby is one New York trial-level decision. It's persuasive and widely covered, but it doesn't bind other courts or set national policy.
  • Private universities aren't bound by constitutional due process the way public institutions are. Their obligation generally runs to their own published policies — which is exactly why Adelphi's failure to follow its own Code of Conduct mattered.
  • Detector accuracy figures move fast. The 2023 Stanford numbers describe the detectors of that period; vendors have since revised their models, and the 2026 VUB results show wide variance between tools. Cite the study and its date, not a number stripped of context.
  • No process evidence, no strong case. If you drafted in one sitting with no version history, this gets harder. Prior writing samples and a substantive conversation about your argument are what's left. Going forward, drafting in a tool that keeps version history is cheap insurance.

FAQ

Can an AI detector be wrong? Routinely. The 2023 Stanford Patterns study found an average 61.3% false-positive rate across seven detectors on human-written TOEFL essays by non-native English speakers. [4] In the Adelphi case, one detector reported 100% AI on a paper two others read as human. [2][3]

Can I be expelled based only on an AI detector score? Policies differ, but a score alone is a weak foundation — and the researchers behind the 2026 VUB study explicitly frame detector output as a signal to look closer, not as proof. [5] The Adelphi court annulled a violation where contrary evidence was ignored and the institution didn't follow its own procedures. [1]

What's the strongest evidence that I wrote something myself? Timestamped version history from Google Docs or Word, showing the document being built incrementally. [6][7] Everything else supports it.

Should I run my paper through a humanizer to prove it's human? No. It doesn't prove authorship, it breaks the match between your text and your version history, and it looks like tampering with disputed work. We make one of these tools and we're telling you not to use it here.

Does Turnitin publish a false-positive rate? Turnitin has reported a document-level false-positive rate of 0.51% and has said the tool is tuned to let roughly 15% of AI writing pass in order to keep false accusations rare. Vendor-reported figures come from the vendor — useful, but not independent.

Do I have a right to see the evidence against me? At a public institution in the US, due process entitles you to notice of the charge, an explanation of the evidence, and an opportunity to respond; FERPA supports access to your own education records. [6] At a private institution, your rights generally come from the school's published policies.

Does it matter which detector flagged me? Yes, materially. The 2026 VUB study found large performance differences across four widely-used tools. [5] Ask which tool produced the score and what its published error rate is.

Sources

  1. Matter of Newby v. Adelphi University, 2026 NY Slip Op 26021 (N.Y. Sup. Ct., Feb. 26, 2026) — law.justia.com. The primary court record. Cited for the holding, the annulment and expungement order, and the court's specific procedural findings.
  2. "When AI Detectors Fail: The Adelphi Plagiarism Case" — studentdisciplinedefense.com, and "Court Rules That University Must Expunge Record…" — lcwlegal.com. Two independent legal analyses of the same ruling; used for the case facts and the court's quoted language, from firms that practice in this area.
  3. "Adelphi student Orion Newby sues over AI plagiarism accusation and wins" — CBS News New York. Established journalism, cited for the case facts and the detail that other detectors read the paper as human-written.
  4. **Liang et al., "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023** — the peer-reviewed source for the 61.3% average false-positive rate and the 97.8% figure. Cited because it's the primary study, not a summary of one.
  5. **"Who wrote this? Evaluating the reliability of AI detection tools in higher education," International Journal for Educational Integrity (Springer), June 2026** — Vrije Universiteit Brussel. Peer-reviewed, current, and the source for the four-detector comparison, the 160-paper methodology, the 1,163-thesis application, and the researchers' "signaling function" conclusion.
  6. GradPilot, "Falsely Accused of AI Cheating? Your Rights and Steps" — cited for the due-process and FERPA framing and the pattern of evidence in cleared appeals. Secondary, used only where it aligns with the primary sources above.
  7. GPTZero, "Falsely Accused of AI Cheating? How to Prove You Didn't" — a detector vendor's own guidance. Included deliberately: when the companies selling detection also tell you process evidence outweighs their scores, that's worth knowing. Note the obvious conflict of interest in both directions — theirs and ours.

Try it on a paragraph — 500 words a day free, no account needed.

Open the editor