Reference file

Voice fingerprint

voice-fingerprint.md

Voice fingerprint

The fingerprint is a short document describing how one author actually writes. It outranks the tell list on every conflict. Build it once per author, then load it before any de-slop pass.

Why it exists

The tell inventory describes the average machine draft. It does not describe a person. Some authors genuinely open with rhetorical questions, lean on parallelism, or use em dashes constantly. Applying the inventory to those authors sands off the habits their audience recognizes them by, and produces a second kind of fake.

Without a fingerprint, a de-slop pass converges every author toward the same neutral register. That register is itself detectable.

Source material

Collect 5 to 10 genuine pieces by the author, in the channel being edited. Genuine means written before they used AI assistance, or verified by them as unassisted.

Mixing assisted and unassisted samples corrupts the fingerprint. The habits picked up from a model get recorded as voice and then defended on every future pass.

If unassisted samples do not exist, note that in the document rather than guessing. A fingerprint built on assisted work is worse than none.

What to record

Sentence length. Shortest, longest, and typical. Note whether they fragment, and how often.

Punctuation. Count em dashes, semicolons, colons, ellipses, and parentheses per thousand words. This is the fastest disagreement to resolve and the most commonly guessed wrong. An author who uses three em dashes per post should keep them.

Openers. How do the last 10 pieces start? Question, claim, number, anecdote, or scene. Record the distribution, not the mode.

Closers. Same. Note whether they land on a next action, a callback, or nothing.

Vocabulary. Words they use that a model would not, and words a model would use that they never do. Both lists are useful. Include their profanity tolerance and their formality level.

Structural habits. Do they use bullets? Headings? Bold? Do they run parallel structures on purpose? Record the ones that show up in most samples.

Contractions. Which ones, and whether the rate changes by channel.

Recurring references. Their standard analogies, the examples they return to, the people and companies they name.

Format

Keep it under 400 words. A fingerprint long enough to be skimmed rather than read stops getting loaded.

Record each item as an observation with a count, not as a rule. "Uses em dashes twice per post on average across 8 samples" is checkable. "Likes em dashes" is not.

Applying it

Load the fingerprint before the first pass, not after. A pass run without it and then corrected against it leaves damage, because the rewrites are already in the author's absence.

On conflict, the fingerprint wins. Every time.

When a flagged construction appears in the fingerprint, leave it and move on. Do not compensate elsewhere, and do not reduce its frequency toward the inventory's preference. Partial deference produces copy that is neither the author nor the neutral register.

Maintenance

Rebuild the fingerprint when the author changes channel or role, and at least once a year. Voice drifts, and a stale fingerprint defends habits the author has dropped.

Flag drift when a new genuine sample conflicts with the recorded fingerprint on two or more items. That is a signal to rebuild, not to override the sample.