AI detectors flag human writing constantly, and here is why

26 August 2026 · 10 min read

Detectors do not detect authorship. They measure how predictable your word choices are and how much your sentence lengths vary. Plain, well-structured, professionally edited writing scores as predictable, which is why it flags. Seven detectors ran over 91 TOEFL essays written by non-native...

A client ran your 2,000-word draft through GPTZero, got back 89% AI, and forwarded the screenshot to your project manager with three question marks. You wrote every word. Now prove it.

Detectors do not detect authorship. They measure how predictable your word choices are and how much your sentence lengths vary. Plain, well-structured, professionally edited writing scores as predictable, which is why it flags. Seven detectors ran over 91 TOEFL essays written by non-native English speakers and wrongly flagged 61.3% of them on average, measured by Liang and colleagues in Patterns in 2023. A score is a probability, not evidence, and you should never accept one as proof.

What a detector score actually measures

Detectors read two numbers off your text. Neither one has anything to do with who typed it.

Perplexity measures how surprised a language model is by your next word. Write "the cat sat on the mat" and a model predicts every word easily, so perplexity runs low. Write "the cat sat on the tax return" and perplexity spikes. Low perplexity reads as machine generated, because models pick likely words by design.

Burstiness measures how much your sentence lengths vary. Take the standard deviation of your sentence lengths and divide it by the mean. Model output lands near 0.30, since almost every sentence falls between 14 and 18 words. Human writing usually runs above 0.60, because people follow a long sentence with a short one without thinking about it.

burstiness 0.08 · mean 15.8 words
Each bar is one sentence, scaled to the longest. Unedited output holds a straight edge because almost every sentence lands between 14 and 18 words. Human writing does not, and the ragged edge is what a detector measures as burstiness. Above 0.60 reads as human.

Both numbers describe the text. Neither describes the writer. A careful human who writes short, plain, evenly paced sentences produces exactly the profile a detector was built to catch, and no amount of honesty on your part changes the arithmetic, because the tool never had access to the question you think it is answering.

The 61.3% problem

The most damaging study on this landed in 2023, when Liang and colleagues at Stanford00130-7) ran seven commercial detectors over 91 TOEFL essays and 88 essays by US eighth graders. The average false-positive rate on the TOEFL essays was 61.3%. On the eighth graders it was close to zero. Same detectors, same day, and the only thing that changed was who wrote the text.

The reason is mechanical rather than malicious. Someone writing in a second language draws on a smaller working vocabulary and leans on constructions they trust. That produces lower perplexity. The detector reads low perplexity, returns a high AI probability, and a human being gets accused of cheating for writing carefully in their fourth language.

The same team proved the mechanism in one move. They fed the TOEFL essays back through ChatGPT with an instruction to use richer vocabulary, then re-ran the detectors. The false- positive rate fell from 61.3% to 11.6%. Making the writing *more* machine-assisted made it look *more* human to the machines, which tells you exactly what these tools measure and how little it has to do with who typed the words.

"Delve" makes the same point from a different angle. Paul Graham said in 2024 that an email containing the word would make him suspect a model wrote it, and the backlash came fast from Nigerian and other African writers, for whom "delve" is ordinary formal English taught in school. The word carries a measured overuse rate of 25.2 times its expected frequency, from Kobak and colleagues in Science Advances in 2025. Both things are true at once. The statistic holds across a corpus and tells you nothing about the individual in front of you.

Scale matters here too. The Pew Research Center checked roughly half a million English web pages in August 2026 and found about 10% carrying signs of AI authorship, and more than a third among pages published since ChatGPT launched. A detector tuned to flag that much of the open web will flag a lot of people who wrote their own sentences.

Six kinds of human writing that flag every time

Our catalogue tracks known false-positive populations alongside the markers themselves, because a detector without a guard list is a liability. Six groups get hit hardest.

Non-native English speakers. The 61.3% figure above. The single largest and most serious failure class.

Formulaic genres written by humans. Legal boilerplate, lab reports, scripture, technical documentation and famous historical texts all run low perplexity by nature. GPTZero returning an AI verdict on the US Constitution became a running joke in 2023 for exactly this reason. The text is predictable because it has been read a billion times.

Professional copyeditors. Editing raises predictability. Someone who removes redundancy, standardises punctuation and evens out paragraph length has moved the text toward the machine profile deliberately, because that is what good editing does.

Neurodivergent and highly structured writers. Consistent sentence patterns are a feature of how many people write, not a sign of automation.

Pre-AI internet genres. LinkedIn broetry dates to around 2017. SEO listicles and corporate style guides predate ChatGPT by a decade. All three flag heavily, and all three were human habits first.

Anyone who uses em dashes. Em dashes were a reliable marker until OpenAI suppressed them in November 2025. Detectors still weighting them are measuring a model generation that no longer exists, and catching the professional writers who have punctuated that way their whole careers.

Every phrase in any slop library was human before it was slop. These are over-production markers measured against a baseline, not inherently bad language.

Why your own editing makes it worse

Clean up your draft and the score gets worse. Almost nobody sees that coming, which is why the writers who care most about their prose tend to be the ones who get accused.

Tighten your prose and you cut the odd word, the tangent, the sentence that ran too long. Perplexity drops. Standardise your paragraphs to a consistent length and burstiness drops with it. Both moves push you toward the profile detectors were built to flag, which means the more professional your final pass, the more machine-like your numbers look.

Writers who paste into a detector, panic at the score, and start swapping words usually make it worse a second time. Word swaps do nothing for rhythm. Sentence length carries more weight in most scoring models than vocabulary does, and it is the one thing almost nobody edits, because you cannot hear your own cadence on the fourth read of your own paragraph.

The way out is not weirder vocabulary. Never insert typos, odd synonyms or broken grammar to move a score, since that damages the writing and fails the only test that matters, which is whether a human editor thinks it reads well. Add specifics instead. A sentence carrying a number, a name or a date raises perplexity honestly, because a model with nothing to say cannot produce one.

What to send back when a client accuses you

Answer within a few hours, and answer with process rather than protest. Protest reads as guilt. Saying "I promise I wrote it" only invites a second opinion from a second detector, and now two numbers are sitting in the thread instead of one.

Send the version history. Google Docs keeps a full revision trail under File, then Version history, then See version history. A document that grew over six sessions across four days does not look like a paste. Screenshot the timeline or share the doc with view access.

Run their detector on something they know is human. Take 500 words from the client's own website, from a page written before 2022, and run it through the same tool. Send both screenshots together. A 70% AI score on their founder's 2019 About page ends the argument faster than any explanation you could write.

Name the false-positive rate. Point at the Stanford 2023 finding00130-7), then at Vanderbilt University, which switched Turnitin's AI detector off on 16 August 2023 and published its arithmetic for doing so. At Turnitin's own stated 1% false-positive rate, roughly 750 of the 75,000 papers Vanderbilt submitted the previous year would have been wrongly flagged. Mention OpenAI as well. The company withdrew its own AI Text Classifier on 20 July 2023, citing a low rate of accuracy, and the published evaluation is worse than most people assume: the tool caught 26% of AI-written text while flagging human writing as machine written 9% of the time. The people who build the models could not build a reliable detector for them.

Offer a standard going forward. Propose that flagged drafts get reviewed against documented markers rather than a probability score, and that you will run that check before delivery. That turns an accusation into a workflow, which is the outcome you want.

Try it. Run the draft they flagged through a marker-based check first. Five checks a day cost nothing and need no signup, and the flag list gives you something concrete to send back instead of a denial.

Check the markers, not the probability

A probability score tells you nothing you can act on. 89% AI does not name a sentence, so it cannot be argued with or fixed. It just sits there.

See which specific phrases a detector is reacting to, and why each one carries weight. The free AI detector matches your text against 908 catalogued terms, 29 structural patterns and 13 formatting tells, then reports every hit with its reason and its reliability tier. The 908 words AI overuses lists the worst of them with the published rate beside each one. Tier one markers are near conclusive on their own. Tier four markers only count in clusters.

Read the structural half in the readability checker. Sentence variance, repeated openers and sticky sentences each get a panel with the offending sentences listed underneath, which is where most accused drafts actually lose.

No language model runs anywhere in that process. The engine matches against a catalogue and applies deterministic rules, so your client's draft never leaves for a model provider. That matters more than it sounds when the work sits under an NDA.

Spelling, grammar and readability run unlimited and free against a 49,000-word dictionary and 16 grammar rules. Unlimited rewrites and the voice cloner sit on Creator at $9.99 a month (about €9.20). Bulk CSV, seats and the API sit on Agency at $29 a month (about €26.70), which is the tier that makes sense once you are checking 40 drafts a month across eight clients.

Euro figures converted at the rate on the published date and rounded. Check the pricing page for the current charge in your currency.

Where detectors are right and this article is wrong

Detectors deserve a fair hearing, so here is the other side.

Detectors catch unedited output reliably. Paste a ChatGPT response straight into a document and most tools will flag it correctly. The failure cases cluster around edited text, hybrid documents and non-native writing. Somebody submitting raw model output and claiming otherwise usually does get caught, and that is the tool working as intended.

Our own markers produce false positives too. A catalogue of 908 terms will flag a statistician who writes "robust" and a watchmaker who writes "meticulous". Density is the guard, not absence. Under two flagged terms per 1,000 words reads as normal variation, and any tool treating a single hit as a verdict has the same problem we are describing.

Hybrid documents beat everyone. A draft where a human wrote the argument and a model smoothed the prose is the hardest class for every detector tested, ours included. No tool on the market resolves that case, and anyone claiming otherwise is overselling.

A clean check is not a guaranteed pass. Our engine removes documented markers and shows you what changed. It cannot promise a GPTZero, Originality or Turnitin result, because those tools weight signals we do not control and update without notice.

Do this before your next delivery

Take the last piece you delivered and run it through the client's detector rather than yours. Note the score. Then run 500 words from the client's own pre-2022 website through the same tool and note that score too.

Keep both screenshots. The gap between them is the argument you will need one day, and gathering it while nobody is accusing you of anything takes about four minutes.

Which client in your roster is most likely to run that check first, and would your last draft for them survive it?

Common questions

Why do vendors advertise a 1% error rate
Most published rates describe whole documents, where a tool only commits above a high confidence bar. Sentence-level rates run several times higher, and a single flagged paragraph is usually what starts an argument. Ask any vendor which of the two numbers they are quoting.
Which detector is the most accurate
None of them well enough to settle a dispute. Accuracy also shifts every time a vendor retrains or a model provider changes its output habits, so a comparison published last year describes tools that no longer behave that way. Pick one, learn its baseline, and never treat it as an arbiter.
What score counts as too high
No threshold is safe, because vendors calibrate differently and change without notice. Compare instead. Run the same detector over text you know is human, from the same author and subject, and read your score against that baseline rather than against 100.
What evidence works if I did not draft in Google Docs
Anything timestamped and messy. Outline files, research notes, voice memos, interview recordings, browser history from the sources you cited, and Slack messages where you argued about the angle. Sequence beats polish, since the point is showing the thinking that came before the prose.
Will a humanizer tool clear a false positive
Sometimes, and it treats the wrong problem. Most humanizers paraphrase through another model, which swaps vocabulary while leaving your sentence rhythm untouched. You end up with a rewritten draft that still scores badly and no longer sounds like you.
Can I ask a client to stop using detectors
Yes, and frame it as a standard rather than a refusal. Offer something concrete in its place, such as version history on delivery plus a marker-based check you run before you send. Clients want assurance, not that particular tool.

Check a draft against all of this

Spelling, grammar, readability and AI phrasing in one pass. No account needed to see what fires.

Run the free check