The words AI overuses — and why generic banlists don't work

Every AI-words-to-avoid list circulating online has the same problem: it was derived from somebody else's writing. Here is what happens when you derive one from your own.

You have seen the list. Delve. Leverage. Robust. Tapestry. Navigate. Elevate. Seamless. Testament. It circulates every few months with a headline about spotting AI writing, and every word on it is genuinely overused.

It is still close to useless as a tool, for reasons we can show with numbers rather than assert.

We built a system that derives these lists empirically instead of borrowing them — it generates model output on your topics in your format, then diffs that vocabulary against your actual writing. What comes out is not quite what the shared lists contain, and the differences are the interesting part.

What the diff actually produced

Run against a corpus of short business posts, the top of the model's over-used list looked like this:

WordRatio (model : human)Model usesHuman uses
innovation10.1×411
below8.4×170
name7.9×160
business6.4×261
challenges5.9×120
whether5.4×110
believe4.9×100
excellence4.4×90
paced3.5×70

Some of those are obvious. Innovation and excellence are on every shared list and deserve to be.

The instructive ones are the entries you would never think to write down.

paced is not a word anyone overuses on its own. It is the back half of a fixed phrase, and pulling the samples confirms it: "In today's fast-paced business environment" opens four different posts, nearly verbatim. The tell is not a word, it is a sentence the model reaches for whenever it has to start something.

below is the same shape from the other end: "Share your thoughts below." "We'd love to hear your thoughts in the comments below." The model has a default closing move, and it deploys it regardless of whether the platform even has comments.

name is the funniest and the most damning. It is not a word at all — it is the model writing "At [Company Name], we believe...", leaving a template placeholder in finished copy. Sixteen times.

whether marks a structural habit rather than a vocabulary one: "Whether it's long-haul trucking, last-mile delivery, or specialized freight..." The three-item conditional opener, once per post, forever.

A word-level list catches innovation. It does not catch "in today's fast-paced business environment," the placeholder, or the compulsive tricolon, because those are not words. This is the first reason generic lists underperform: the strongest tells are phrases and structures, and they only surface when you diff against a real person writing about the same things.

Why a borrowed list is worse than no list

Three problems, all measured.

It doesn't know your vocabulary. On a small corpus of 2,709 words, 16 of 41 automatically derived entries turned out to be noise. Eight of them appeared in the writing samples included in the same prompt — the model was shown a paragraph containing the word share while being told never to write share. Another eight were fragments of hashtags, already covered by a separate rule. One banned word appeared in the user's own topic description. A borrowed list has this problem permanently and invisibly, because it was derived from a corpus that is not yours.

It confuses topic with style. Ratio ranking will happily flag shipping, supply chain, and commerce as machine tells when the corpus is about logistics. They are not tells. They are the subject. This is why the diff has to run on your own topics — compare your essays about databases against model text about hiking and you have measured the topic, not the voice.

It bans things that make writing more human. Contractions are the clearest case. It's, don't, and we're show up as statistical differences constantly, and banning them produces text that is measurably further from a human voice, not closer. Our extractor refuses to put a contraction on the list for that reason. A generic list has no such guard.

And banning words does not work anyway

This is the part that undercuts the entire premise of a banlist as a prompt instruction.

Adding "NEVER use these words" to the prompt, with a correct, empirically derived list attached, moved violations from 8.84 to 8.78 per thousand words. Eight generations. A 0.7% difference, which is far inside run-to-run noise.

What did change was which words appeared. Before the instruction, the hits were ahead, business, company, leadership, means, progress, share. After it: ahead, growth, means, share, shipping, success. The model dropped several words we had named and substituted neighbors we had not. It complied locally and routed around the constraint globally.

That is the expected behavior, not a bug. Examples and instructions in a prompt apply pressure toward things; they cannot reliably apply pressure away from a high-probability word, because there are always four more words nearby with nearly the same probability. Ban six and you get the seventh.

Detecting the words in the finished draft and rewriting those spans moved the same measure to 0.0, every run.

The list is not the mechanism. The checker is the mechanism. A list you can enforce in code is worth more than a longer list in a prompt.

The half nobody publishes

Every AI-words list is a list of commissions — things the model adds. The omissions matter at least as much and get no attention.

Frontier models underuse colloquialisms, hedges, and discourse glue: the small connective words that carry no information and all of the personality. So the same diff that produces the over-used list also produces the under-used one, and that second list is what keeps the output from going sterile.

Strip the slop without restoring the glue and you have not made the text sound like you. You have made it sound like nobody — clean, competent, characterless. Which is arguably a worse outcome, because generic-but-fluent is harder to notice and just as unconvincing to a customer.

How to derive one for yourself

The procedure is not complicated, and the order matters.

  1. Collect real writing in the format you will generate. Review replies if you are generating review replies; not your About page. Roughly 15,000 words before a banlist means anything — below that, absence is chance rather than evidence. Format statistics like sentence length stabilize much earlier, around 3,000 to 5,000 words.
  2. Generate model output on topics pulled from that same corpus, in the same format and length. Anything else measures subject matter.
  3. Diff the two vocabularies by rate rather than raw count, with a floor on occurrences so a single stray word never earns a place on the list.
  4. Read the result before you trust it. Delete anything that appears in your own writing, anything that is topic rather than style, and every contraction.
  5. Keep the under-used list too, and require those words to appear rather than merely permitting them.
  6. Enforce it with a checker, not a sentence in a prompt. This is the only step that changes the output.

Step four is the one people skip, and it is the reason to distrust any tool that hands you a banlist without showing you the derivation.

Where this lands for review replies

A business replying to its own Google reviews has a corpus of maybe four hundred words — forty short thank-yous written over two years. That is nowhere near enough to claim a word is absent from someone's voice, so the honest thing to do at that size is skip the banlist entirely, build the format statistics that are stable at that volume, and collect more writing before claiming anything stronger.

Any vendor who tells you they have modeled your voice from thirty past replies is describing a list they got from somewhere else.


The mechanism behind all of this is covered in more depth in why AI can't write like you. If you want it running on your review page rather than explained, that's the service.