Your client highlighted one word in a 1,400-word draft and asked whether you wrote it. The word was "delve". That question is how a $2,000 monthly retainer starts to die, and answering it badly costs you the account.
What one flagged word costs a working writer
Run the math on a flagged draft. A ghostwriter charging $400 per post loses roughly six hours to the rewrite, the defensive email and the follow-up call. That drops the effective rate from $66 an hour to about $40. Lose the client outright and the number goes to zero.
Agency writers carry a worse version of the same risk. A three-person shop shipping 40 posts a month across eight clients only needs one account manager to paste a draft into a detector and screenshot the result. The conversation that follows is never about the writing. It is about whether the retainer renews.
The frustrating part is that most flagged drafts were written by a person. Someone used a model for the outline, kept a few of its phrases, and never noticed that four of them sit in the top 50 most over-represented terms in published corpus research. The work was real. The vocabulary gave it a signature it did not earn.
Where the overuse numbers actually come from
Vocabulary advice usually arrives as somebody's hunch. These numbers do not.
Kobak and colleagues published the largest of the studies in Science Advances in 2025, comparing 15 million PubMed abstracts from 2010 to 2024. They put a floor on it too: at least 13.5% of 2024 abstracts show signs of having been through a model, which is roughly 200,000 papers in a single year. "Delves" turned up at 25.2 times its expected rate. "Underscores" reached 9.1 and "showcasing" 9.2.
Liang and colleagues, presenting at ICML in 2024, ran the same approach across ICLR 2024 peer reviews. "Meticulous" appeared at 34.7 times its baseline rate, the highest single-word multiplier in that dataset. "Intricate" reached 11.2 times. "Commendable" reached 9.8.
Detector vendors published their own frequency work on top of that, and it deserves a different level of trust. Pangram measured "vibrant tapestry" at roughly 17,000 times its expected rate and "serves as a testament" at around 4,000. GPTZero put "today's fast-paced world" at 107 times. IsGPT logged "left an indelible mark" at 317 and "an unwavering commitment" at 202.
Treat those as the vendors' own claims rather than as findings. Each company sells a detector, none of the numbers has been replicated by anyone else, and none comes with a published method you could check. They point the right way. The peer-reviewed figures are the ones to quote when a client asks you to prove something.
Juzek and Ward added a useful wrinkle at COLING in 2025. The overuse is not random drift. It traces back to preference tuning, which means these words got selected for on purpose because human raters liked them. Your model reaches for "commendable" because somebody once clicked the thumbs-up.
None of this is confined to academia. The Pew Research Center ran roughly half a million English web pages through a detection model in August 2026 and found 10% of them showing signs of AI authorship, rising to more than a third of pages published after ChatGPT launched. The same study clocked em dashes at double their 2023 rate, and negative parallelism, the binary contrast construction, at nearly triple.
Our catalogue compiles all of it into 908 terms across 18 categories, alongside 29 structural patterns, 13 formatting tells and 12 statistical features. Every entry carries a reliability tier, because a term measured at 17,000 times baseline deserves more weight than one measured at 9.
The 33 phrases with measured overuse rates
Start with the phrases that carry a published multiplier. These are the ones where the evidence is strongest and the fix is cheapest. Paste a draft into the free AI detector and every one of them that fires is listed with its reason.
| Phrase | Overuse rate | Source | Write instead |
|---|---|---|---|
| vibrant tapestry | 17,000× | Pangram | name the actual mix |
| in the ever-evolving | 11,000× | Pangram | delete the clause |
| serves as a testament | 4,000× | Pangram | shows, proves |
| important to note | 3,000× | Pangram | delete it, then state the point |
| provide a valuable insight | 468× | IsGPT | say what the insight was |
| left an indelible mark | 317× | IsGPT | changed, shaped |
| an unwavering commitment | 202× | IsGPT | state what they did |
| a stark reminder | 166× | IsGPT | delete, keep the fact |
| gain a comprehensive understanding | 120× | IsGPT | learn, work out |
| a nuanced understanding | 115× | IsGPT | name the distinction |
| today's fast-paced world | 107× | GPTZero | cut the opener entirely |
| aims to explore | 50× | GPTZero | this post covers |
| meticulous | 34.7× | Liang, ICML 2024 | careful, thorough |
| delves | 25.2× | Kobak, Sci Adv 2025 | digs into, gets into |
| showcasing | 9.2× | Kobak, Sci Adv 2025 | shows, puts on display |
| underscores | 9.1× | Kobak, Sci Adv 2025 | shows, points to |
| intricate | 11.2× | Liang | complicated, fiddly |
| commendable | 9.8× | Liang | good, solid |
Multipliers show how much more often the phrase appears in model output than in comparable human text. Rounded from the published figures.
Two entries in that list deserve special attention. "Important to note" and "today's fast-paced world" both open sentences, which means a detector sees them in a structural position as well as a lexical one. Openers cost you twice.
Density beats presence, and that trips up most writers
The single biggest mistake writers make with a word list is treating it as forbidden vocabulary. That reading produces stilted drafts and does nothing for the score.
"Robust" is the right word in a statistics paper. "Meticulous" is the right word for a watchmaker. One flagged term in 1,200 words is noise, and any detector calibrated sensibly ignores it.
Six flagged terms in one paragraph is a model. That is the actual signal, and it is why our catalogue scores cumulatively across independent families rather than counting hits. A draft picks up weight when lexical tells, structural patterns and formatting habits all point the same way at once.
Work to a rate, not a rule. Aim for under two flagged terms per 1,000 words, and treat any paragraph carrying three or more as a rewrite rather than a find-and-replace job. Replacement alone tends to shuffle the problem sideways. A sentence built around "an unwavering commitment to excellence" has no content in it, and swapping the phrase for "a strong commitment to excellence" leaves the emptiness exactly where it was.
The 18 categories your draft falls into
Term lists usually arrive as one long alphabetical dump, which makes them useless at the keyboard. Ours splits into 18 groups because different groups fail in different places.
Verbs, 148 terms. The largest group and the one that shows up first. "delve", "unlock", "foster", "harness", "leverage", "elevate", "empower", "streamline".
Adjectives, 89 terms. "robust", "seamless", "comprehensive", "transformative", "pivotal", "invaluable", "unparalleled".
Nouns, 66 terms. "tapestry", "testament", "beacon", "realm", "cornerstone", "landscape", "ecosystem". Metaphor props, mostly, standing in for a concrete noun the writer never chose.
Openers, 33 terms. The costliest group per word, because openers repeat. A draft where four paragraphs start the same way reads as generated even when every individual word is clean.
Closers, 20 terms. "in conclusion", "at the end of the day", "the bottom line".
Transitions, 39 terms. "additionally", "furthermore", "moreover", "conversely".
Hedges and editorializing, 40 terms. "arguably", "it is worth noting", "some might say".
Puffery phrases, 48 terms. The empty-claim family. "a game-changing approach to."
Engagement bait, 49 terms. The LinkedIn tell. "Thoughts?" "Agree?" "Drop a comment below."
Chat artifacts, 28 terms. Text that leaked straight out of the interface. "Certainly!" "I hope this helps." "As an AI language model", which Pangram clocked at 294,000 times baseline and which still turns up in published posts every week.
The remaining groups cover academic phrasing, adverb intensifiers, email outreach patterns, quantifier vagueness, fiction and roleplay habits, model signature names and the 33 documented multipliers above. Fiction carries 78 terms of its own, largely because character names give models away. "Elara" appears in model-written fiction at roughly 85,000 times its baseline rate.
Fix the six worst offenders in your next draft
Skip the full list on a deadline. These six carry the most weight for the least work.
One. Cut every opener that scene-sets. Delete "in today's fast-paced world", "in an era where", "when it comes to". Open on the fact instead. A draft loses nothing and picks up a stronger first line.
Two. Replace "delve" with "dig into" or "get into". Do the same for "unlock", "harness", "leverage" and "foster". Plain verbs, every time.
Three. Delete the phrase "it is important to note". Whatever follows it is either important, in which case say it, or it is not, in which case cut it. Pangram measures that formula at roughly 3,000 times its expected rate, which makes it one of the loudest single signals in the whole catalogue.
Four. Kill the metaphor nouns. "tapestry", "testament", "beacon", "landscape", "realm". Each one is standing in for something specific the draft never named. Name it.
Five. Check every list of three. Models bundle into triads whether or not three things exist. Read the third item. If it restates the second, delete it. Two specific items beat three padded ones.
Six. Delete every "it's not just X, it's Y". Both halves make the same claim. State what the thing is.
Work through those six on a 2,000-word draft and expect about 25 minutes the first time, under ten once the patterns stick.
Words are only half the tell
Vocabulary is the part writers fix, which is exactly why it stopped being the part that catches them.
Structure is the other half, and our catalogue tracks 29 structural patterns against 908 lexical ones. Sentence-length variance sits at the top. Unedited model output runs a standard deviation over mean near 0.30, because sentences cluster between 14 and 18 words. Human writing sits above 0.60. A writer who edits vocabulary and leaves rhythm untouched keeps the loudest signal in the draft.
Formatting adds 13 more. Bold lead-ins on every bullet. Emoji as section markers. A colon before every list. Headings that all start with a gerund.
Punctuation moves faster than the rest, so treat any punctuation rule as temporary. Em dashes were a reliable marker until OpenAI suppressed them in November 2025. Anyone still running an em-dash-based check is measuring last year's models and flagging a lot of professional writers who have used em dashes their whole careers.
Read one paragraph of your draft aloud and count the syllables in each sentence. Sentences that all land the same length are the problem, and no word list will surface that for you. The readability checker reports the variance directly, and AI detectors flag human writing constantly covers what happens when that number gets you accused of something you did not do.
Run the check before you send, not after the client replies
Catch the flagged terms while the draft is still yours to fix. Paste it in, read the slop score, and work down the flag list from tier one.
The score counts across all six families at once, so a draft carrying two lexical hits and a repeated-opener pattern rates worse than a draft with four lexical hits and clean structure. That weighting matters, since the second draft is the one a human editor actually reads as fine.
Open the Editing report for the structural half. Sentence variance, repeated openers and sticky sentences each get their own panel, with the offending sentences listed underneath.
The engine runs no language model. It matches against the catalogue, applies 285 word replacements and 141 phrase replacements, and reports every change it made. Your draft never goes to a model provider, which matters when the text belongs to a client under NDA.
Try it. Five de-slops a day cost nothing and need no signup. Paste a real client draft and read the flag list before you decide anything.
Spelling, grammar and readability run unlimited on the free tier against a 49,000-word dictionary and 16 grammar rules. Unlimited rewrites and the voice cloner sit on Creator at $9.99 a month (about €9.20). Bulk CSV, seats and the API sit on Agency at $29 a month (about €26.70).
Euro figures converted at the rate on the published date and rounded. Check the pricing page for the current charge in your currency.
Where this advice breaks down
Three honest limits, because a word list oversold does more damage than no word list.
Detectors flag human writers constantly. Stanford researchers reported in 2023 that detectors misclassified 61% of TOEFL essays written by non-native English speakers as AI generated. Formal registers, neurodivergent writing patterns and anyone trained in academic prose all sit in the same failure zone. A flag is evidence of a pattern, not proof of authorship, and any writer accused on a detector score alone should say so.
Tells decay. The em dash stopped working as a marker within weeks of a single model update. Every list on this page, ours included, describes the models of the last two years. Treat any specific term as perishable and the underlying habit, which is empty sentences wearing expensive vocabulary, as permanent.
A clean score is not a bypass. Our engine removes documented markers and shows you what changed. It cannot promise a GPTZero result, an Originality result or a Turnitin result, because those tools weight signals we do not control and update without notice. Anyone selling you a guaranteed pass is selling something they cannot deliver.
The workaround for all three is the same and it is not a tool. Put a checkable fact in every paragraph. A sentence carrying a number, a name or a date reads as human because a model with nothing to say cannot produce one.
Start here tomorrow
Pull up the last draft you sent a client. Search it for six strings: "delve", "testament", "tapestry", "it is important to note", "in today's", and "unwavering". Count the hits.
Under two across the whole piece and your vocabulary is fine, so go straight to sentence length instead. Three or more and you have a paragraph to rewrite rather than a word to replace.
Which piece in your queue this week would survive that search, and which one would you rather fix before the client runs it?