Common AI Words and Phrases: An Evidence-Ranked List
Every AI word list is an assertion. This one is a measurement. We computed the excess frequency ratio for all 407 style words in the source dataset behind the leading study, and most of the words on popular lists do not survive it.
"Delves" is the most documented AI word in existence: it appeared 28.2 times more often in 2024 academic writing than the pre-ChatGPT trend predicted. But it's an outlier. We computed the excess frequency ratio for all 407 AI style words in the dataset behind the leading peer-reviewed study, and only 23 of them clear 5x. Ninety show effectively no signal at all, and 47 of the most-quoted AI words on the internet, including "tapestry" and "treasure trove," do not appear in the research anywhere. The full ranked list, the method, and its limits are below.
What You'll Learn
- The 23 AI words with the strongest measured evidence behind them, with exact ratios
- How the excess frequency ratio is calculated, and how to reproduce it yourself
- Why "delves" carries four times the signal of "delve," and what that means for editing
- Which words on every popular blacklist the data does not support
- Why models overuse these words, according to the study that tested the mechanism
- Which tells are expiring as the words leak into ordinary human speech
- What this analysis cannot tell you, stated plainly
- A copy-paste prompt that flags by evidence strength instead of running a flat blacklist
What Are AI Words and Phrases?
AI words and phrases are terms that large language models produce at a measurably higher rate than the pre-2022 human baseline, which makes their presence a statistical signal that a passage was drafted or edited by a model.
The word "signal" is doing the work in that sentence. None of these words are AI-only. Every one existed in English before ChatGPT, and every one is still used by human writers every day. What changed is the rate.
The useful unit is the excess frequency ratio, written as r. It answers a single question: how many times more often did this word appear in 2024 than it should have, if the trend from before ChatGPT had simply continued? An r of 1.0 means the word is exactly on trend and tells you nothing. An r of 28 means something happened.
Most AI word lists skip this step entirely. They publish a set of words with no number attached, which means you have no way to tell a 28x signal from a 1.01x coincidence. Both look identical on a bulleted list.
What Are the Strongest AI Words, Ranked by Evidence?
These are the 23 style words with an excess frequency ratio of 5.0 or higher. Every figure is computed from the open dataset published alongside Kobak et al., Science Advances, July 2025. The count column is the number of 2024 PubMed abstracts containing the word.
Two families dominate this table. Four of the top twelve are forms of "delve." Five of the top twenty-three are forms of "underscore." If you only remember two roots from this entire article, remember those.
The Next Tier: Clear Signal, 3x to 5x
Fifty-three words sit here. Treat one instance as unremarkable and three in the same paragraph as a pattern.
intricately (5.0x), encompassed (4.9x), groundbreaking (4.9x), encompassing (4.7x), emphasizing (4.7x), overlooking (4.4x), consolidates (4.3x), offering (4.2x), necessitating (4.2x), substantiates (4.2x), aligning (4.1x), formidable (4.1x), revolutionizing (4.1x), emphasising (4.0x), showcases (3.8x), meticulous (3.8x), swift (3.8x), surpasses (3.8x), showcased (3.8x), adept (3.8x), heightened (3.8x), encompass (3.8x), bolstering (3.8x), comprehending (3.7x), endeavors (3.6x), advancements (3.6x), uncharted (3.6x), poised (3.6x), emerges (3.5x), unveiled (3.4x), notable (3.4x), stands (3.4x), unraveling (3.3x), urging (3.3x), aiding (3.2x), scrutinizing (3.2x), hinting (3.2x), burgeoning (3.2x), holds (3.2x), pinpointed (3.2x), realms (3.2x), fostering (3.1x), advocating (3.1x), discernible (3.1x), pivotal (3.1x), seamlessly (3.1x), leveraging (3.1x), detrimentally (3.1x), capitalizing (3.0x), enhancements (3.0x), dependable (3.0x), fosters (3.0x), equipping (3.0x)
Moderate Signal, 2x to 3x
A further 122 words land here, including several that popular lists treat as smoking guns: transformative (3.0x), multifaceted (2.9x), exceptional (2.8x), nuanced (2.7x), noteworthy (2.7x), interplay (2.5x), leverages (2.4x), discern (2.3x), streamline (2.3x), nuances (2.3x), crucial (2.1x), insights (2.0x), valuable (2.0x).
A 2x ratio is real but modest. On its own it should not change an editorial decision.
The complete ranked list of all 407 words, with ratios, signal bands, 2022 and 2024 frequencies, and raw abstract counts, is available as a CSV and a JSON file at the end of this article.
How Was This List Calculated?
The method is the study's, not ours. We applied it to the study's own published data and extended it to every word in the set.
- Source data. Kobak, González-Márquez, Horvát, and Lause analyzed every PubMed abstract from 2010 through 2024 and published two files openly: a yearly occurrence matrix of 362,442 words across 15 years, and an annotated list of 900 excess words. Our totals reconstruct to 15,103,887 abstracts, matching the 15.1 million reported in the paper.
- Filter to style words. The authors annotated each of the 900 words as style, content, or other. Content words are domain vocabulary that spiked for real-world reasons, like "alphafold" or "omicron." Artifacts like "armonk" and "aspx" are marked other. Only the 407 words annotated style are AI writing tells. Every other list in this category ignores this distinction.
- Build the counterfactual. For each word, frequency is the share of that year's abstracts containing it. The expected 2024 frequency is a linear extrapolation from the 2021 and 2022 values, which is the last clean read before LLM assistance became widespread.
- Compute the ratio. Excess frequency ratio r is the observed 2024 frequency divided by the expected 2024 frequency. For very high-frequency words we also report the excess gap δ, the absolute difference in the share of abstracts.
How Do We Know the Method Is Right?
Because it reproduces the numbers the authors published. The paper reports six figures. Ours match all six.
That validation is the reason to trust the other 401 numbers. The method is not our invention and the data is not our data. The only thing we added is the arithmetic nobody had run.
Why Do Most AI Words Turn Out to Be Weak Signals?
Because the distribution is far more lopsided than any list implies. Here is how all 407 style words actually break down.
Just 6% of documented AI style words carry strong evidence. More than half sit below 2x, and 90 of them show essentially nothing.
This is the practical finding buried in the research. The category people treat as one flat blacklist is a small hard core wrapped in a very large fuzzy edge. Editing against the whole set means spending most of your attention on words that were never telling you anything.
Which AI Words Does the Data Not Support?
Two separate groups fail, and they fail for different reasons.
Group One: In the Research, But Near Zero Signal
These words appear in the excess vocabulary set and are frequently cited as AI tells. Their ratios say otherwise.
A ratio below 1.0 means the word became less common than trend predicted. "Verifies" and "realizes" moved in the opposite direction from the one the folklore claims.
"However" is the one worth sitting with. It gets flagged constantly. Its ratio is 1.01, which is another way of saying nothing happened to it at all.
Group Two: Not in the Research Anywhere
These 47 words are among the most-quoted AI tells online. None of them appear in the 900-word excess vocabulary set:
amplify, beacon, boast, boasts, characterized, cognizant, complementary, complexity, conceptualize, critique, deep, dive, embark, empower, endeavor, enlightening, facilitate, foster, furthermore, garner, holistic, integral, journey, kaleidoscope, leverage, moreover, navigate, offerings, paramount, pave, pertinent, profound, recognize, relentless, resonate, robust, synergy, systemic, tapestry, testament, treasure, trove, underpinnings, unlock, unravel, vibrant
That includes "tapestry" and "treasure trove," the two words most people would name first if you asked them for an AI word.
The honest caveat, stated up front: this does not prove those words are innocent. See the limitations section below. It proves something narrower and still useful, which is that no published frequency data currently supports them.
Why Does AI Overuse These Words?
The intuitive answer is that models absorbed these words from their training data. Two researchers at Florida State University tested that and found the mechanism sits somewhere else.
Tom Juzek and Zina Ward identified 21 words with anomalous frequency spikes in scientific abstracts, then ran an online study mimicking corporate reinforcement learning from human feedback. Participants rated pairs of abstracts, one carrying buzzwords like "delve," "realm," and "intricate," one without. Separately, a Llama variant trained on human preference data was less surprised by buzzword-heavy abstracts than the base model was.
Their conclusion, published at COLING in January 2025, is that human preference data is a source of these lexical preferences. The models didn't learn "delve" from books. They learned it from people rating their output.
Is the Nigerian English Theory True?
In April 2024, Alex Hern proposed in The Guardian's TechScape newsletter that "delve" traces to RLHF annotation work outsourced to Nigeria, where the word appears more often in formal business English. The theory spread quickly, helped along by Paul Graham publicly calling "delve" a marker of AI writing and Nigerian writers pushing back hard.
Here is the accurate version, and nearly every article on this topic skips it. Juzek and Ward's peer-reviewed work supports the mechanism, that human preference data drives the overrepresentation. It does not confirm the origin story about Nigerian annotators specifically. One is demonstrated. The other is a plausible hypothesis that has not been tested.
Repeating the origin story as settled fact is the single most common error in this category.
Why Does "Delves" Score Four Times Higher Than "Delve"?
Because inflected forms carry the signal, and no blacklist reproduces this.
The same pattern holds for other roots. "Underscores" is 13.8x while "underscore" is 6.7x. "Meticulously" is 11.3x while "meticulous" is 3.8x. "Surpassing" is 7.1x while "surpass" is 1.85x, which drops it out of the meaningful range entirely.
The practical consequence: a find-and-replace on the root word is the wrong operation. "Surpass" is fine. "Surpassing" is not. If your editing checklist says "delve" and stops there, it's tuned to the weakest member of its own family.
Are These AI Tells Expiring?
Some are, and this is the part every static list gets wrong.
Juzek, Anderson, and Galpin analyzed 22.1 million words of unscripted spoken English, including conversational podcasts, and presented the results at AIES in October 2025. Nearly three-quarters of the AI-associated words they tracked rose in spoken use after ChatGPT's release. Some more than doubled. Confirmed risers included delve, intricate, surpass, boast, meticulous, strategically, and garner.
The detail that makes it persuasive: "underscore" rose considerably while its synonym "accentuate" did not. People aren't simply speaking more formally. They're absorbing the specific words the models favor.
Cross-referencing that list against our ratios shows the decay is uneven:
The words at the bottom of that table are the ones to stop flagging first. "Strategically" and "surpass" already sat below 2x, and they are actively diffusing into ordinary speech. Their remaining diagnostic value is close to zero.
This is also why we date this page and recompute it rather than treating the list as fixed. A word list without a timestamp is a claim about a moving target.
What Are the Most Common AI Phrases?
Phrases are a different problem. A word is a frequency question. A phrase is usually a structural habit wearing vocabulary as a costume.
None of these can be fixed by find-and-replace. "It is important to note that X" doesn't become human when you delete the opener, because the problem is that the sentence was built to signal importance rather than demonstrate it. You have to rewrite the claim, not the wrapper.
Wikipedia's WikiProject AI Cleanup maintains the most complete catalog of this kind of pattern. We broke down Wikipedia's Signs of AI Writing list separately, including a prompt you can paste into your editor.
What Should You Write Instead?
The pattern across every row: the AI version names a category of importance, the human version names the actual thing.
What Tells Does a Word List Miss?
Most of them, and this is the honest limit of a page like this one.
Vocabulary is the shallowest layer. Underneath sit patterns no word list catches:
- Uniform sentence length. Human writing varies. Model output tends toward an even rhythm, sentence after sentence landing in the same 15 to 22 word band.
- Paragraph symmetry. Three paragraphs of three sentences each, repeated down the page.
- Rule of three by default. Lists of three items where the content justified two or five.
- Balanced hedging. Every claim followed by its own counterweight, so nothing is ever actually asserted.
- Formatting residue. Bold-colon lead-ins on every bullet, title case in headings, stray markdown, emoji in a professional document.
- Confident vagueness. Fluent prose that survives deletion of any given sentence without losing meaning, because no sentence carried a specific fact.
You can score zero on every word in the table above and still write something that reads as machine-made, because the tell was never the vocabulary. That's the core argument of our guide to humanizing AI text for marketing content, and it's why word swapping is the last step of the job rather than the first.
Do These Words Actually Get Content Flagged?
Not directly, and the evidence points the other way.
AI detectors do not run a banned word list. They score statistical properties of text, mainly perplexity, which measures how predictable each token is, and burstiness, which measures variation in sentence length and structure. A passage full of "delves" and "pivotal" that varies its rhythm can score as human. A passage with none of these words and perfectly even sentence structure can get flagged.
Detector accuracy is also worse than the marketing implies. Originality.AI cites University of Washington work finding that even trained human raters identify AI text at roughly 55%, which they fairly describe as basically a coin flip.
There's a fairness problem underneath this. False positives land hardest on non-native English speakers, whose formal register overlaps with model output, and on anyone writing in an academic register. The Paul Graham episode is the public version of a thing that happens quietly to students and freelancers constantly.
That's the practical case against blacklists. If you edit to beat a word list, you've optimized against the wrong target and possibly made your writing worse. We covered what actually triggers a flag in why detectors say you used AI when you didn't.
What Are the Limits of This Analysis?
Stating these plainly is what separates a reference from a listicle. Four limitations matter.
The corpus is biomedical. Every figure comes from PubMed abstracts. Words that are rare in biomedical writing are underrepresented no matter how common they are in AI-written marketing copy. This is the main reason "tapestry" does not appear: biomedical researchers had little occasion to write it before or after ChatGPT. Absence from this dataset is weak evidence, not proof of innocence.
The register is academic. Abstracts are formal, structured, and heavily edited. Tells that show up in conversational or marketing prose may not surface here at all.
The measurement window ends in 2024. Models released since then have their own habits, and some 2024 tells have decayed. The AIES 2025 spoken-English finding is direct evidence that the ground is moving.
The counterfactual is an extrapolation. Projecting the 2024 expectation from 2021 and 2022 assumes those years were themselves clean and that trends were roughly linear. It's the authors' own method and it reproduces their published figures, but it is a model, not a measurement.
What survives all four caveats: for the specific words in the table, in formal written English, the ratios are real, reproducible, and far better evidence than anything else currently published on this topic.
The Prompt: Flag by Evidence, Not by Blacklist
Paste this into Claude, ChatGPT, Gemini, or Perplexity. It flags by signal strength and never auto-replaces, because swapping flagged words for synonyms produces hedged, thesaurus-driven prose that reads worse than the original.
You are an editor checking a draft for AI writing tells. You work from evidence,
not folklore. Flag, never auto-replace.
STEP 1: FLAG BY SIGNAL STRENGTH
The number after each word is its excess frequency ratio: how many times more
often it appeared in 2024 writing than the pre-ChatGPT trend predicted.
STRONG (flag every instance, 5x or higher):
delves 28x, underscores 14x, delved 12x, meticulously 11x, showcasing 11x,
intricacies 10x, expediting 9x, delve 8x, intricate 8x, underscoring 7x,
surpassing 7x, delving 7x, commendable 7x, underscore 7x, excels 6x,
pioneers 6x, grappling 6x, renowned 6x, realm 5x, garnered 5x,
revolutionize 5x, escalating 5x, underscored 5x
CLEAR (flag if two or more appear in one paragraph, 3x to 5x):
intricately, encompassed, groundbreaking, encompassing, emphasizing, overlooking,
consolidates, offering, necessitating, substantiates, aligning, formidable,
revolutionizing, showcases, meticulous, swift, surpasses, showcased, adept,
heightened, encompass, bolstering, comprehending, endeavors, advancements,
uncharted, poised, emerges, unveiled, notable, stands, unraveling, urging,
aiding, scrutinizing, hinting, burgeoning, holds, pinpointed, realms, fostering,
advocating, discernible, pivotal, seamlessly, leveraging, capitalizing,
enhancements, dependable, fosters, equipping
MODERATE (note only if the draft is dense with them, 2x to 3x):
pinpointing, transformative, advancing, aligns, unparalleled, multifaceted,
crafting, inquiries, uphold, accentuates, exceptional, harnesses, akin, nuanced,
elevates, interplay, leverages, paving, discern, elucidates, streamline, nuances,
empowers, crucial, insights, valuable, noteworthy, illuminates, scrutinize
DO NOT FLAG (below 2x, the evidence does not support treating these as tells):
however, significant, analysis, using, based, during, between, this, were,
comprehensive, enhance, elucidate, harness, facilitates, landscape, imperative,
strategically, elevate, unlocking, invaluable, surpass, additionally
STEP 2: FLAG STRUCTURAL TELLS
- Sentence length variance. If most sentences land between 15 and 22 words, say so.
- Paragraph symmetry. Repeated blocks of the same sentence count.
- Rule of three used by default, where the content justified two or five.
- Negative parallelism. "Not only X, but also Y."
- Balanced hedging. Every claim counterweighted so nothing is asserted.
- Vague attribution. "Studies show," "experts agree," with no named source.
- Formatting residue. Bold-colon bullets, title case headings, stray markdown, emoji.
- Confident vagueness. Any paragraph you could delete without losing a specific fact.
STEP 3: REPORT
Return a table: quoted phrase | tell type | signal strength | why it reads as AI.
Then answer one question plainly: which paragraphs contain a specific fact that
only this author could know, and which contain none?
Do not rewrite anything. Do not suggest synonyms. Flag and explain only.
HumanizeAI Framework References
This article is an application of the E in H.E.A.R.T., Evidence Over Claims, taken further than usual. H.E.A.R.T. requires that every factual assertion be traceable. A word list is nothing but factual assertions, one per word, which means an unsourced list is a page composed entirely of unbacked claims. Ranking by measured ratio is what the principle looks like when you follow it all the way down.
It also demonstrates the second of the GEO Visibility Framework's Three Citation Tests: does AI have proof? A page listing 74 words with no sourcing gives an answer engine nothing safe to attribute. A page stating that "delves" carries an excess frequency ratio of 28.2, computed from a named open dataset using a method that reproduces the authors' published figures, gives it a claim it can quote and stand behind. That difference is most of what separates a page that gets cited from a page that gets scraped.
Founder Observation
I stopped keeping a banned words list about a year into running HumanizeAI, and it wasn't a principled decision. The list stopped working.
We'd strip every flagged word out of a draft, read it back, and it still sounded like a machine wrote it. Clean vocabulary, same dead rhythm. Meanwhile writers on the team started avoiding perfectly good words because they'd seen them on a list somewhere, and the writing got worse in a new way: hedged, thesaurus-y, visibly working around something.
Running the actual numbers for this article was uncomfortable in a specific way. Ninety of the 407 words show essentially no signal. "However" is 1.01x. We had been editing against noise and calling it rigor.
What changed our output was a different question. Instead of asking which words to remove, we started asking which sentence in this paragraph contains a fact only we could know. Usually the answer was none. That's the real tell, and no word list will ever catch it, because the problem was never that the model chose "pivotal." The problem is that the paragraph didn't know anything.
[FOUNDER OBSERVATION: Steve, this is written to the angle you picked and the data backs the middle section now. To publish, it needs one concrete anchor. The specific draft or client project where the cleaned-up version still read as AI, roughly when it happened, or the writer who told you the list was making their work worse. One real detail. I have not invented one.]
Research and Supporting Evidence
- Kobak, D., González-Márquez, R., Horvát, E.-Á., and Lause, J. "Delving into LLM-assisted writing in biomedical publications through excess vocabulary." Science Advances, Vol. 11, No. 27, eadt3813, 2 July 2025. Analyzed over 15 million PubMed abstracts from 2010 to 2024 and concluded that at least 13.5% of 2024 abstracts were processed with LLMs, reaching 40% in some subcorpora, an impact on scientific vocabulary exceeding that of the Covid pandemic. arXiv:2406.07016 | Science Advances
- Kobak et al. open dataset. berenslab/llm-excess-vocab. Contains results/excess_words.csv, the 900 excess words with author annotations, and results/yearly-counts.csv.gz, a 362,442 by 15 yearly occurrence matrix. Updated July 2025. All ratios in this article were computed from these two files. GitHub
- Juzek, T. and Ward, Z. "Why Does ChatGPT 'Delve' So Much? Exploring the Sources of Lexical Overrepresentation in Large Language Models." Proceedings of the 31st International Conference on Computational Linguistics (COLING), January 2025. Identified 21 words with anomalous frequency spikes and concluded through RLHF-mimicking experiments that human preference data, not training corpora, is a source of these lexical preferences. FSU release, 17 February 2025 | ACL Anthology
- Juzek, T., Anderson, B., and Galpin, R. "Model Misalignment and Language Change: Traces of AI-Associated Language in Unscripted Spoken English." Proceedings of the Eighth AAAI/ACM Conference on AI, Ethics, and Society (AIES), October 2025. Across 22.1 million words of unscripted spoken English, nearly three-quarters of tracked AI-associated words rose in use after ChatGPT's release, with some more than doubling. FSU release, 26 August 2025
- Gray, A. "ChatGPT 'contamination': estimating the prevalence of LLMs in the scholarly literature." arXiv:2403.16887, 25 March 2024. Estimated that at minimum 60,000 papers, slightly over 1% of all 2023 publications, showed LLM assistance based on disproportionate increases in marker keywords including meticulous, commendable, and intricate. arXiv:2403.16887
- Gillham, J. "Most Commonly Used ChatGPT Words and Phrases." Originality.AI, 17 October 2025. Analysis of 12,517,475 words of ChatGPT output, cleaned to 10,784,010, found delve at 146 occurrences, embark at 139, nuances at 109, beacon at 85, imperative at 85, and endeavor at 82, concluding that many widely publicized marker words are not in fact common in model output. originality.ai
- Hern, A. "TechScape: How cheap, outsourced labour in Africa shaped AI English." The Guardian, 16 April 2024. Origin of the widely repeated Nigerian English hypothesis for "delve," presented in the original as a hypothesis. theguardian.com
- Wikipedia:Signs of AI writing. WikiProject AI Cleanup, continuously updated. The reference taxonomy for structural and formatting tells rather than vocabulary. Wikipedia
Mini Case Study: What the Ratios Change in Practice
Illustrative composite, built from the ratios above.
Take a paragraph that a conventional blacklist would light up:
"This comprehensive analysis delves into the intricate landscape of customer retention. However, it is important to note that the findings underscore a significant opportunity to leverage existing data."
A flat blacklist flags eight words here: comprehensive, delves, intricate, landscape, however, findings, underscore, leverage.
The ratios say four of those eight are noise. "However" is 1.01x. "Analysis" is 1.01x. "Landscape" is 1.47x. "Comprehensive" is 1.76x. Editing them out costs you time and buys you nothing.
The four that matter are "delves" at 28.2x, "underscore" at 6.7x, "intricate" at 7.8x, and "leveraging" at 3.1x. That's a much smaller, much more actionable edit.
But the real problem with the paragraph isn't on either list. It contains no fact. Strip every flagged word and you still have a sentence that says a study looked at a topic and found something. That's the tell a word list structurally cannot see.
Key Takeaways
- Only 23 of 407 documented AI style words carry an excess frequency ratio of 5x or higher. Ninety show effectively no signal.
- "Delves" is the strongest documented AI word, appearing 28.2 times more often in 2024 than pre-ChatGPT trends predicted.
- Inflected forms carry the signal. "Delves" is 28.2x while "delve" is 7.9x, and "surpassing" is 7.1x while "surpass" is 1.85x.
- Forty-seven widely cited AI words, including tapestry, treasure trove, and testament, do not appear in the peer-reviewed excess vocabulary research at all.
- "However" has an excess ratio of 1.01, which means the word is exactly on trend and carries no diagnostic value.
- The cause appears to be RLHF human preference data rather than training corpora, per Juzek and Ward's COLING 2025 study.
- Some tells are expiring. Delve, intricate, meticulous, strategically, surpass, and garner are all rising in unscripted human speech.
- AI detectors score perplexity and burstiness, not vocabulary, so editing against a word list optimizes against the wrong target.
- The tells a word list cannot catch are structural: uniform sentence length, symmetric paragraphs, default rule of three, and fluent prose containing no specific fact.
Frequently Asked Questions
What are the most common AI words?
The AI words with the strongest measured evidence are delves, underscores, delved, meticulously, showcasing, intricacies, expediting, delve, intricate, and underscoring. Each appeared at least seven times more often in 2024 academic writing than pre-ChatGPT trends predicted, based on an analysis of 15.1 million PubMed abstracts published in Science Advances in July 2025.
What is the single most common AI word?
"Delves" carries the highest documented excess frequency ratio at 28.2, meaning it appeared roughly 28 times more often in 2024 than the pre-ChatGPT trend predicted. Its other forms also rank highly: delved at 12.3, delve at 7.9, and delving at 7.0. In absolute terms the word remains uncommon, appearing in 5,152 of roughly 1.44 million 2024 abstracts.
Is "tapestry" an AI word?
There is no published frequency evidence that it is. "Tapestry" does not appear in the 900-word excess vocabulary set identified by Kobak et al. from 15.1 million biomedical abstracts. That corpus is academic and biomedical, so a word rare in that register would be underrepresented regardless, which means absence is not proof. What can be said accurately is that no peer-reviewed frequency data currently supports treating "tapestry" as an AI tell.
Are em dashes really an AI tell?
Em dashes are a weak signal at best. They appear frequently in unedited model output, which is why Wikipedia's WikiProject AI Cleanup lists them among formatting patterns to watch, but they are also standard punctuation used by professional human writers for over a century. An em dash alone proves nothing. An em dash combined with uniform sentence length, symmetric paragraphs, and no specific facts is a stronger combined signal.
Do AI words actually get content flagged by AI detectors?
No, not directly. AI detectors score statistical properties of text, primarily perplexity, which measures how predictable each word is, and burstiness, which measures variation in sentence length and structure. They do not run a banned word list. Text containing many so-called AI words can pass as human if its rhythm varies, and text containing none can be flagged if its structure is uniform.
Is it wrong to use these words if I wrote the text myself?
No. Every word on every AI word list is a legitimate English word that predates large language models. Avoiding a word purely because a list mentions it usually makes writing worse, producing hedged or thesaurus-driven prose. The useful question is whether the word is the most precise available choice, not whether it appears on a list.
Do AI words change as models update?
Yes, in both directions. New models develop new lexical habits, and older tells weaken as words diffuse into human writing. A 2025 study of 22.1 million words of unscripted spoken English found that nearly three-quarters of tracked AI-associated words increased in human speech after ChatGPT's release, including delve, intricate, meticulous, and garner. Any AI word list should be treated as dated rather than permanent.
Why does ChatGPT use "delve" so much?
Research from Florida State University points to reinforcement learning from human feedback rather than training data. Tom Juzek and Zina Ward found that a model variant trained on human preference data was less surprised by buzzword-heavy text than the base model, concluding that human preference data is a source of these lexical preferences. A separate theory attributing "delve" to Nigerian English usage among outsourced annotators is widely repeated but has not been demonstrated.
Will Google penalize content that uses AI words?
No. Google's stated position is that it rewards helpful, original content regardless of how it was produced, and there is no vocabulary-based penalty. Content that reads as generic can underperform, but the cause is thin or unoriginal substance rather than the presence of specific words.
Can humans reliably tell AI writing from human writing?
Not reliably. Originality.AI cites University of Washington research finding that even after training, human raters identified AI-generated content at roughly 55% accuracy, close to chance. This is why frequency evidence and structural analysis are more useful than intuition, and why accusations based on a reader's gut feeling about word choice are unreliable.
What is an excess frequency ratio?
An excess frequency ratio compares how often a word actually appeared in a given year against how often it should have appeared if the pre-ChatGPT trend had continued. A ratio of 1.0 means the word is exactly on trend and carries no signal. A ratio of 28 means the word appeared 28 times more often than expected. The measure comes from Kobak et al.'s 2025 Science Advances study of excess vocabulary in biomedical abstracts.
Download the Full Dataset
All 407 style words with excess ratios, signal bands, 2022 and 2024 frequencies, expected frequencies, and raw abstract counts.
- humanizeai-ai-style-words-2026.csv (spreadsheet)
- humanizeai-ai-style-words.json (machine-readable)
Free, no signup. Attribution to this page appreciated if you use it.
Try It On Your Own Draft
Word lists tell you what to remove. They don't tell you what's missing, which is usually the real problem.
HumanizeAI rewrites AI-drafted content for rhythm, specificity, and voice rather than swapping words against a blacklist, and the AI Article Agent builds drafts that carry evidence from the start. Run a paragraph through and compare it against the structural tells above.
Additional Resources
- How to Humanize AI Text: The Complete Guide for Marketers (pillar)
- AEO and GEO: The Marketer's Complete Guide to AI Search Visibility (pillar)
- Wikipedia's Signs of AI Writing: The Full List, Plus a Copy-Paste Prompt
- Why Does ZeroGPT Say I Used AI When I Didn't?
- HumanizeAI AI Article Agent
About the Author
Steve Palomares is the founder of HumanizeAI, a Paloma Digital LLC product, where he builds tools and frameworks for content teams working on AI search visibility. He is the author of the H.E.A.R.T. writing framework and the GEO Visibility Framework, and he writes about AEO, GEO, and content authority from the operator's seat rather than the analyst's. More about Steve
Analysis computed 18 August 2026 from the Kobak et al. open dataset. Last updated: 18 August 2026.