Why AI can't write like you — and what actually fixes it

You gave it five samples of your writing. It came back ninety percent right, and then one word ruined it. Here is the mechanism behind that word, and the only thing we found that fixes it.

Your AI writes with an accent.

You have probably already run the experiment. You pasted in five samples of your own writing, asked for something in the same voice, and got back a draft that was ninety percent right. The structure was yours. The register was yours. Then one word landed that you have never typed in your life. Robust. Leverage. Delve. Comprehensive. You deleted it, moved on, and quietly stopped trusting the tool for anything that goes out under your name.

That one word is not a small failure of an otherwise working system. It is the system working exactly as designed, and understanding why changes what you build to fix it.

Persona is not voice

Two different things get called "tone" and they behave nothing alike.

Persona is role and stance: what the speaker knows, how they treat the reader, whether they are warm or clipped, whether they apologize or explain. Persona responds well to instructions. "You are a patient studio owner who thanks people by name" works. You can write that in a sentence and the model will hold it for a whole conversation.

Idiolect is the token-level distribution underneath: which specific words you reach for at a choice point, how long your clauses run, where your punctuation clusters, which hedges you use, what your rhythm does across a paragraph. It is the linguistic fingerprint that makes a text message from your sister recognizable before you see the name.

Persona is prompt-tractable. Idiolect is not, past a fairly low ceiling.

This matters because the stray word is almost never a persona failure. The model got the role right. It was warm, it was grateful, it named the customer. What leaked through was idiolect — a word from its own distribution rather than yours, in a slot your samples never demonstrated.

Why the word leaks

Here is the part that most explanations skip.

Style examples in a prompt apply positive pressure only. They pull the output toward the patterns they show. What they cannot do is apply targeted negative pressure, because absence is not a signal the model reasons over. It does not look at fifteen thousand words of your writing and conclude, "she never once wrote robust, therefore suppress robust." Nothing in the mechanism performs that inference. Your corpus is evidence of what you do, and silent about everything you don't.

So at every word, the model blends a soft pull toward your demonstrated style against a very strong prior built from its pretraining and its alignment training. On common words the two agree and nothing interesting happens. But at a choice point — a content word with four or five viable synonyms, in a sentence shape your samples never covered — the prior usually wins. One word later the model is back in agreement with your style, and the sentence continues correctly.

That is the whole explanation for why it is one or two words and not a whole paragraph. The leak is local because the disagreement is local.

The accent analogy is exact. Someone imitating a regional accent can learn the vocabulary and the famous vowels. What gives them away is the word they didn't rehearse, where the native pronunciation asserts itself for one syllable before the imitation resumes.

What we measured

We built a voice-fidelity system to test this properly, with a deliberate stage structure so each lever could be measured separately instead of assumed. One stage adds a corpus in context. One adds measured statistics. One adds an empirically derived list of words to avoid. One detects violations in code after generation and rewrites only those spans.

Two of those stages are directly comparable, and the comparison is the most useful number we have.

Adding the instruction "NEVER use these words" to the prompt moved violations from 8.84 to 8.78 per thousand words. That is a 0.7% change across eight generations — which is to say, nothing. Run-to-run noise on the same configuration is larger than that by an order of magnitude.

Detecting the same words in code after generation and rewriting them moved the same measure to 0.0. Not 0.4. Zero, in every run, from baselines of 6.2 and 16.2 in two separate comparisons.

Zero is not a lucky sample. It is a floor, because a deterministic check either finds the word or it doesn't.

Telling the model not to do something does almost nothing. Checking its output and fixing what you find works completely.

There is a stranger detail in that first comparison. The banlist instruction did not reduce the rate of generic words — but it did change which generic words appeared. Before the instruction: ahead, business, company, leadership, means, progress, share. After: ahead, growth, means, share, shipping, success. The model dropped some of the words we named and replaced them with neighbors we hadn't. It complied locally and routed around the constraint globally, which is exactly what a positive-pressure mechanism would be expected to do.

"Just give it more examples" does not work either

This is the first thing everyone tries, and it has been tested more rigorously than our eight generations.

Language models approximate style acceptably in structured formats — news copy, transactional email, anything with strong conventions — and struggle with informal writing, which is where personal voice actually lives. Crucially, adding more demonstrations does not close the gap. More samples buy you register and topic. They do not buy you the distribution.

The failure is also bidirectional, which is the half that gets ignored. Frontier models overuse a recognizable set of vocabulary and prefer standardized grammar and nominalizations, and they underuse colloquialisms, hedges, and discourse glue — the small connective words that make prose sound like a person talking (OpenReview).

That second half is why "strip out the AI words" produces something worse than what you started with. Remove the commission errors without restoring the omissions and you get text that is clean, competent, and sterile. Not your writing. Nobody's writing.

The three things that actually work

Measure, don't describe. "Conversational and direct" is not enforceable — there is nothing in the sentence a checker can act on. "Sentences average ten words with a standard deviation of seven, replies run one line, no contractions" is enforceable, and it is also true in a way an adjective never is. The variance figure matters as much as the mean; consistent sentence length is one of the strongest tells that a machine wrote something.

Give every constraint a detector. This is the measured result from the previous section, generalized into an architecture. Any rule worth having — never quote a price, never name a competitor, never invent a customer name — is worthless as a line in a prompt and reliable as ten lines of code that inspect the finished draft. Instructions are weak. Filters are strong.

Restore the glue, not just remove the slop. Every voice has connective habits: the hedges, the sentence openers, the small words that carry no information and all of the personality. Those have to be identified and required, or you are only doing half the job.

Two ways measurement goes wrong

Having said "measure, don't describe," it is worth showing the two ways we got measurement wrong, because both are easy to repeat and neither is obvious.

A rate across many outputs is not a rule for one output. A profile might record that a business uses 0.56 emoji per post, or includes a call to action 15.6% of the time. Both are accurate descriptions of the body of work. Neither is a constraint you can hand to a generator writing a single post, because you cannot put 0.56 of an emoji in one place. We made this mistake three separate times before naming it, and the failure is asymmetric: told about a 15% rate, the model either applies it to everything or to nothing. A habit that appeared in 12% of the source material became 0% in the output, because the model correctly reads a minority pattern as "don't."

The distinction that fixes it: per-item constraints, like length and whether you open with a greeting, can be enforced on a single output. Distribution constraints, like how often you use an emoji, can only be checked across a batch and have to be handled differently or left alone.

A number that summarizes a structure is not the structure. One profile recorded roughly four exclamation marks per reply, so the constraint became "use about four exclamation marks." The output came back with !! after every single sentence. Arithmetically correct. Structurally nothing like the source, which actually varied: one mark here, three there, a plain sentence in between, in sequences like [3,2], [1,3,2], [2,2,1].

Four marks is 2.1 runs of average length 1.9, and the average erased the only part that mattered. The fix was to measure the runs and their distribution, state the variation as an explicit rule, and back it with a checker. Measure the structure you want reproduced, not a scalar that happens to summarize it.

Voices are real objects, and they are measurable

The strongest evidence that any of this is more than vibes came from running the same pipeline over two yoga studios in the same city, both replying to five-star Google reviews, both warm, both grateful.

Studio AStudio B
Words per reply19 (stdev 3.6)37 (stdev 13.9)
Exclamation marks3.7 per reply1.2 per reply
Run lengths! 47% !! 33% !!! 20%! 100%, never doubled
Sentences with no !6%60%
Greetingnone — straight inown first line, 100% of replies
Emojinone0.9 per reply
Contractions per 100 words01.75
Sentences opening with And/But/So20%0%

Same industry, same city, same star rating, same eight-week window. Two completely different writing systems.

And the detector built from each studio's statistics rejects the other studio's real replies and passes its own. That is the result that settles the question. These are not stylistic preferences we projected onto the data. They are signatures, stable enough that a deterministic checker can tell them apart without a model in the loop.

The honest limit

Everything above has a ceiling, and it is worth stating plainly because most people selling voice-matching will not.

With prompting alone, a judge asked to pick which of two passages a human wrote plateaus somewhere around 0.65 to 0.75 accuracy on personal voice, where 0.50 means indistinguishable. Length is the most stubborn residual: generated replies consistently overshoot the target by 15 to 50% even after a style pass, because verbosity is closer to the model's core than vocabulary is.

Not every layer earns its tokens, either. We tested swapping a static set of writing samples for a retrieval step that picks examples matched to the current topic. When the whole corpus already fit in the prompt, retrieval made things slightly worse — it won 1 comparison out of 8, and the measured distance from the target format went up rather than down. Most of what it retrieved was already in the prompt.

Getting past the prompting ceiling requires changing the weights, and the efficient path is contrastive rather than imitative: pairs of (what the model wrote, what you would have written instead). Fewer than ten such pairs have been shown to beat few-shot prompting, supervised fine-tuning, and self-play by 19 points, because the pair ("leverage" → "use") teaches a boundary that no instruction can express.

That is the real fix, and it starts where the measurement ends.

Why any of this matters for a review reply

A customer reading your response to their review can tell. Not analytically — they are not counting exclamation marks — but the same way they can tell a form letter from a note. Prospective customers scrolling your profile can tell too, and there are more of them than there are reviewers.

A generic reply under a warm review is worse than no reply, because it converts a compliment into evidence that nobody was listening. Getting the voice right is not a finishing touch on review automation. It is the entire thing that makes review automation worth doing.

If you want to see what this looks like running on a real business's review page, that's the service.