We asked AI to reply to reviews for two yoga studios. The replies were measurably different.
Same model, same prompt architecture, same industry, same star rating. The only variable was whose past replies went in — and every measured axis came apart.
"Warm and friendly" describes both of these studios accurately, and it would generate neither of their replies.
That is the short version of what we found when we built voice profiles for two hot yoga studios in the same metro area. Both reply to nearly every five-star review. Both are genuinely grateful in a way you can feel through the screen. Both would describe their own tone with the same three adjectives.
Their actual writing has almost nothing in common.
The setup
Two studios. We'll call them Studio A and Studio B. Same city, same discipline, same format of review coming in — short five-star posts naming an instructor and the room. We took each studio's own past replies and built a measured profile from them, then ran both through an identical pipeline: same model, same prompt structure, same constraints. The only variable was whose writing went in.
Studio A's corpus: 7 review-and-reply pairs, 136 words of replies total. Studio B's: 10 pairs, 400 words.
Those are small numbers and we will come back to what they do and don't support. They were enough.
What came back
| Studio A | Studio B | |
|---|---|---|
| Words per reply | 19 (stdev 3.6) | 37 (stdev 13.9) |
| Words per sentence | 9.1 | 12.5 |
| Exclamation marks per reply | 3.7 | 1.2 |
| Run lengths | ! 47%, !! 33%, !!! 20% | ! 100% — never doubled |
| Replies varying run length internally | 83% | 0% |
| Sentences with no exclamation | 6% | 60% |
| Lines per reply | 1 | 2 |
| Greeting on its own line | 0% | 100% |
| Emoji per reply | 0 | 0.9 (60% of replies) |
| Contractions per 100 words | 0 | 1.75 |
| Sentences opening with And/But/So | 20% | 0% |
| Names the reviewer | 100% | 100% |
| Vocabulary variety (type-token ratio) | 0.54 | 0.30 |
One row is the same and every other row is different.
In prose: Studio A writes a single unbroken line, no greeting, straight into the thanks, exclamation marks stacked and varied — one here, three there, deliberately uneven. No emoji, no contractions, and one sentence in five starts with and or so, which is how people write when they are typing quickly and happily.
Studio B opens with the reviewer's name on its own line, every time. One exclamation mark per reply at most, never doubled. Sixty percent of its sentences have no exclamation at all. An emoji lands in about six replies out of ten. The replies run twice as long and are twice as consistent in structure, with a lower vocabulary variety — the same phrases recur, which is what a house style looks like when more than one person at the front desk is writing.
Here is roughly what each sounds like, with names changed:
Studio A: "Awww thank you Maya!!! So glad to hear you're having an awesome experience with all of our teachers!!"
Studio B: "Hi Maya,
Thank you so much for the great review ⭐️ We are so glad you enjoyed class! We hope to see you on your mat again soon."
A regular customer at either studio would recognize their own studio's reply instantly and find the other one slightly off, without being able to say why.
The result that settles it
Measuring two things and finding different numbers is not, by itself, interesting. Any two writers differ on some axis if you look at enough axes.
The test that matters is whether the profiles are discriminative: build a deterministic checker from Studio A's statistics, feed it Studio B's real replies, and see what happens.
Each studio's checker rejects the other studio's real replies and passes its own.
No model in the loop, no judgment call — a plain program that measures length, run structure, line count, greeting placement and emoji rate against the expected distribution and flags anything outside it. Studio B's genuine, human-written replies fail Studio A's checker. Not because they are bad replies. Because they are somebody else's.
That is the strongest evidence we have that these are real signatures rather than stylistic decoration we projected onto small samples. A vibe cannot be cross-tested. A distribution can.
Why this is the whole argument for voice
Consider what a generic review-reply tool has to work with. It knows the star rating, the review text, the customer's name, and possibly a tone setting: friendly, professional, casual.
Set both studios to "friendly" and you get the same reply for both, which means you get the wrong reply for at least one of them and probably both. The friendly-sounding default a model produces — a greeting line, a measured two sentences, one exclamation mark, a closing invitation — happens to land fairly close to Studio B and nowhere near Studio A. Studio A's actual voice, with its stacked exclamation marks and its no-greeting opening, is something a model will essentially never produce on its own, because it sits well outside the polite-professional default that alignment training pushes toward.
So the tool would quietly rewrite Studio A into a different business. Not badly. Just not them, forever, under every review, on the page their next hundred customers read before booking.
A statistical profile is not a nicer way of saying "brand guidelines." It is the difference between a constraint a program can check and an adjective everybody interprets differently.
What we have not measured
Two limits, stated plainly, because the numbers above are more persuasive than the study design deserves in a couple of places.
The sample sizes are small. Seven pairs and ten pairs. Format statistics like reply length and line count stabilize early and we trust those. Anything requiring a claim about absence — this studio never uses a particular word — needs far more text, on the order of 15,000 words, before absence means anything more than chance. We do not make those claims at this corpus size, and neither should anyone else.
Both corpora are entirely five-star. Every reply either studio has written in the sample is an answer to praise. Which means the profile can tell you exactly how Studio A thanks someone and has nothing whatsoever to say about how Studio A apologizes. A model asked to generate an apology from a gratitude-only profile will improvise, and improvisation is precisely where these systems invent things — a discount nobody authorized, a commitment nobody made.
That is why, in the production version of this, any review carrying a complaint routes to a person rather than to the generator, regardless of how many stars it has. The profile earns trust in the register it was built from and gets no credit outside it.
What to take from it
If you are evaluating anything that writes on your behalf, the question to ask is not whether it sounds good. It is whether the vendor can show you a measurement of your voice, and whether that measurement can tell your writing apart from a competitor's.
If the answer is a paragraph of adjectives, the tool is generating from adjectives, and every business that picked the same three words is getting the same replies you are.
The mechanism behind this — why prompting alone cannot get you there — is in why AI can't write like you. If you want your own profile built and running, start here.