
The most useful thing anyone can learn about talking to and around artificial intelligence right now is this: the flattening isn’t a metaphor. Actual measurable shifts in vocabulary and speech patterns are already detectable in ordinary human conversation, and the mechanism behind it is one linguists have studied for decades.
Key Points
- A 2024 preprint found measurable, population-level shifts in word usage toward ChatGPT-associated vocabulary following the tool’s release — the first empirical evidence of humans imitating an LLM’s language habits.
- The underlying mechanism, known in sociolinguistics as entrainment or accommodation, is not new; what’s new is a nonhuman, mass-scale interlocutor shaping it.
- Deliberately keeping filler words, pauses, and small imperfections is emerging as a practical strategy for sounding — and staying — recognizably human, though rigorous audience-perception data on this specific claim is still thin.
- AI writing and transcription tools tend to smooth away hesitation markers by default, so preserving a personal voice increasingly requires active effort rather than passive habit.
- Speech-recognition systems trained on “normative,” fluent voices already penalize accents and disfluencies, raising the stakes of treating imperfection as noise rather than identity.
How Convergence Actually Works
Long before chatbots existed, linguists documented a phenomenon called entrainment: when two people talk, they gradually converge on shared vocabulary, pacing, and even pitch. It’s why couples start finishing each other’s sentences and why call-center workers unconsciously pick up a caller’s regional cadence. Entrainment is not a flaw in human communication — it’s a feature that builds rapport and mutual understanding. The complication is that entrainment doesn’t care whether the other party is a person. Research on human-computer interaction shows speakers adjust their phrasing and delivery in real time even when they know they’re addressing a machine, adapting to what they assume the system needs to understand them.
That adaptive instinct, harmless when the “conversation partner” is a single voice assistant, becomes something else entirely when hundreds of millions of people are entraining daily to the same handful of large language models. A 2024 study tracking word usage before and after ChatGPT’s public release identified a statistically significant surge in terms disproportionately associated with the model’s own output — the first hard evidence that everyday speech is drifting toward machine-generated norms rather than the reverse. Commentary from linguists elsewhere has pointed to specific words, such as “delve,” spiking in popularity for reasons traceable to the demographics of the workers who helped train these systems — a texture of the model’s training data quietly becoming a texture of ours.
Why Polish Isn’t the Same as Authenticity
The instinct to clean up speech — cut the “ums,” tighten the sentence, eliminate the pause — feels like an improvement. In writing assisted by AI, however, that instinct routinely goes too far. Guidance aimed at researchers and professionals warns that vague prompts like “improve this” or “make it clearer” tend to iron out precisely the constructions that make a piece of writing sound like a specific person rather than a competent generality; the fix recommended is to deliberately reintroduce strategic tentativeness, personal phrasing, and the kind of hedges a careful human actually uses when uncertain. Separate analysis of AI-generated text finds the same pattern from the other direction: machine writing defaults to monotonous, evenly-paced sentence structures, while human writing carries small irregularities — a slightly awkward clause, a sentence that doesn’t quite land “by the book” — that readers register, consciously or not, as evidence of a mind rather than a model.
Spoken language works the same way, and this is where the popular advice to “let your filler words be” earns its credibility. Verbalizing a thought mid-formation — the audible “so, what I’m getting at is” or the pause before a hard word — signals something an optimized script cannot: that a live cognitive process is happening in real time. The weakness in the current evidence base is worth naming honestly. Most of the research supporting voice preservation addresses accent, dialect, and written style; direct experimental proof that audiences specifically judge filler-laden speech as more trustworthy or relatable than polished speech is still sparse. The claim is intuitive and consistent with adjacent findings, but it hasn’t yet been isolated and tested the way the lexical-drift research has.
The Structural Pressure Toward Sameness
What makes this more than an aesthetic preference is who gets erased when “clear” becomes the only acceptable register. Automatic speech recognition systems, trained overwhelmingly on fluent, standard-accent datasets, already fail disproportionately for speakers with accents, stutters, or nonstandard rhythms — not because those speech patterns are less intelligible to other humans, but because the training data never learned to value them. Commentary on AI’s effect on language more broadly warns of a feedback loop: platforms reward standardized, efficient, emotionally flat phrasing; people begin mimicking the models that produce it; and models, in turn, retrain on that increasingly homogenized human output. Accent-preservation researchers frame the countermeasure explicitly as conservation work — treating regional speech patterns the way a museum treats an artifact at risk of disappearing, worth recording and actively defended rather than assumed to persist on its own.
None of this means fluency is bad or that AI tools are the enemy of clear communication — for many contexts, from technical documentation to safety instructions, standardized clarity genuinely serves the audience better than idiosyncrasy. The distinction that matters is between deploying AI as an editor that flags unclear passages for a human to fix in their own words, versus letting it draft and finalize prose that then becomes the default voice heard, imitated, and normalized. The Transmitter’s advice to “use AI as an analyzer, not a writer” captures the operative principle: keep the human in the loop as the source of idiom, hesitation, and idiosyncrasy, and let the machine do inventory rather than composition.
What Practical Vigilance Looks Like
Sounding human in the AI era, then, is less about performing imperfection than about refusing to let a tool quietly delete it. That means resisting the reflex to smooth every pause out of a recorded talk, declining to let a transcription service “clean up” a quote until the speaker’s actual cadence disappears, and noticing when a draft — yours or a chatbot’s — has become suspiciously even. The open empirical question is how far this goes: whether audiences reliably prefer disfluent, personal speech to polished alternatives across contexts like leadership communication, teaching, or customer service remains understudied, and the honest answer is that fillers and pauses could just as easily read as uncertainty in a job interview as they read as authenticity in a podcast. Context still governs judgment. What the current evidence supports without much argument is the underlying trend: left unchecked, convergence toward machine-shaped language is already measurable, and the burden of preserving a distinct voice now sits more with the speaker than with the tool.
Sources:
time.com, accentify.co.uk, linkedin.com, ie.edu, diva-portal.org, medium.com, frontiersin.org, community.openai.com













