Why Banning AI Buzzwords Does Not Make AI Text Read Human
AI words-to-avoid lists are collections of phrases that appear more often in AI output than in human writing, which editors remove so drafts stop sounding templated.
On this page
Quick Answer
AI words-to-avoid lists are collections of phrases that appear more often in AI output than in human writing, which editors remove so drafts stop sounding templated. Our July 2026 measurements show that bans work literally: banned words appeared in 0-12 of 816 AI-drafted articles. Bans cut shared phrasing across a fleet. They do not remove the underlying model style: in a separate September test, 81-100% of generated prose sections still read AI to a detector trained on that pipeline. The clearest difference we found came from substance, on one site: sections built on the company's own records read AI 37% of the time, against 71% for the rest.
A list of AI phrases to avoid answers a narrower question than most teams assume. It can stop a fleet of articles from sharing the same wording. On the evidence we have, it does not make a passage read as though a person wrote it.
The shared wording is real and measurable. In July 2026 we measured 816 AI-drafted articles across 24 websites. One stock opening phrase, "the short answer," appeared in 74% of them, on every site. One scripted sentence appeared in 362 articles on 18 sites, and in a 30 July snapshot of 843 articles the most common article skeleton, the same section order, was shared by 95 articles on 11 sites.
Bans fixed what they named. Banned words appeared in 0-12 of the 816 articles, 98-100% compliance. But a ban executes exactly as written: banning that phrase as a heading moved it into the first sentence. And our first per-site phrasing generator, built to vary wording, re-inserted the same phrase into one site's pool until we added a phrase-level guard.
Outside the pipeline, the rules that matter are not written in phrases either. Amazon's KDP guidelines classify a book's text by who created it, and the U.S. Copyright Office analyzes human authorship case by case. Detector accuracy is a separate question: the University of San Diego Legal Research Center's guide, "The Problems with AI Detectors: False Positives and False Negatives," cites studies that found detectors "neither accurate nor reliable." A phrase audit answers none of those questions.
This article separates the two problems: shared phrasing across a fleet, which bans can fix, and the underlying model style, which they do not.
Under U.S. Copyright Office guidance, copyright does not extend to purely AI-generated material, and whether human contributions amount to authorship is analyzed case by case. Amazon KDP requires authors to disclose AI-generated text even after substantial edits. Neither line is drawn at vocabulary.
John Overman's newsletter post "AI Copyright: Lesson Learned" quotes Amazon's KDP content guidelines directly: "If you used an AI-based tool to create the actual content (whether text, images, or translations), it is considered 'AI-generated,' even if you applied substantial edits afterwards." AI-assisted content, by contrast, is content the author created and only edited, refined, error-checked, brainstormed or improved with AI tools, and KDP does not require authors to disclose it. The Copyright Office report he cites adds that using AI "to assist rather than stand in for human creativity" does not affect copyright protection.
An editor who strips flagged phrases from an AI-drafted manuscript has therefore not moved it from "AI-generated" to "AI-assisted" under Amazon's definitions. Phrase removal changes the word inventory, not who created the content.
Detection accuracy adds a separate problem. The University of San Diego guide reports that neurodivergent students and students writing in English as a second language are flagged by AI detectors at higher rates than native English speakers. A Stanford HAI report, AI-Detectors Biased Against Non-Native English Writers, found that seven detectors were near-perfect on essays by U.S.-born eighth-graders but classified 61.22% of TOEFL essays by non-native English students as AI-generated. Phrase bans do nothing about that gap.
Why do AI detectors produce false positives on human writing?
One plausible reason is that neutral institutional prose sits close to the register AI models imitate. Three of the four pre-2021 Medium posts an earlier version of our detector called AI were institutional.
We collected 2,868 Medium posts that Common Crawl archived before 31 December 2020, one post per author. An earlier version of our detector called 4 of them AI (0.14%). Three were institutional prose: a US government aid programme update, a policy analysis of an Affordable Care Act bill and a corporate design case study. The fourth was a list of 100 headline-style post titles. There were no false positives in tech tutorials, data science, crypto and finance, startups, personal essays or culture writing. The current version called 2 of the 2,868.
Government writing tested the pattern. On 637 explanatory documents from US federal agencies archived before 2021, the previous version called 8 AI (1.26%), above our 0.5% cap, and plain-English copy written for international readers was the hardest group. After we added 100 documents from other agencies to training, the new version called 2 of the 637, and 0 of 511 documents from 8 agencies it had never seen. Across 9,963 verified human documents in total it called 4 AI, and on a 4,546-document pool it called none (95% upper bound 0.08%).
A banned-words list does not address what those misfires had in common, which was register rather than vocabulary. Whether stripping buzzwords pushes a draft further toward that neutral register is a plausible concern that we have not measured.
Style tells behave the same way. In one company's AI-drafted articles, a sanitizer had replaced every em dash with a spaced hyphen, and the text still carried 12.9 dash-punctuation marks per 1,000 words against 0.8 in its own pre-AI blog. Swapping the character kept the rhythm. A budget of 2 dash marks per 1,000 words and rewritten asides brought new articles to 1.69.
Our detector gives two separate readings that are never blended: origin, meaning who most likely wrote the text, and style, meaning how generic or distinctive it reads. A human writer can produce generic prose, which is exactly the case our false-positive cap exists for: at most 0.5% of human documents called AI, in every register.
Outside tools show a wide spread. The University of San Diego guide reports that Turnitin cites a false-positive rate below 1%, while a Washington Post study found 50% for one detector, though with a much smaller sample. In AI detectors rarely flag human writing, but sometimes miss AI text (Epoch AI, July 2026), Pangram and GPTZero flagged 0 of 495 human passages and Originality.ai flagged 19.
Evasion can be cheap. Cat Casey, chief growth officer at Reveal and a member of the New York State Bar AI Task Force, is quoted in the guide saying she can fool detectors 80-90% of the time by adding the single word "cheeky" to her prompt. Turnitin says its checker can miss roughly 15% of AI-generated text and that it accepts this because it does not want to flag human-written text as AI. The company frames that as a deliberate tradeoff between the two errors, and no phrase ban changes it.
What makes some article section types read AI 100% of the time while others read AI only 8%?
In our data the split followed how much the model wrote from scratch: decision frameworks read AI 100% of the time, while author bios built from a real profile read AI only 8%.
We scored 2,625 visible sections from 146 AI-drafted articles written for 13 websites in September 2026, each section separately. Decision frameworks read AI in all 16 cases, outlook charts 99% of the time, before-and-after comparisons 94% and myth-vs-fact boxes 87%. Calls to action read AI 36% of the time and human 33%; curated resource lists 42% and 21%; author bios 8% and 51%. As a control, human Medium posts cut into 50-, 100-, 200- and 300-word chunks were called AI 1.0%, 0%, 0% and 0% of the time, so short sections were not flagged just for being short.
Shared phrasing was not what separated them. Phrases shared across different websites made up 0.5% or less of every prose section, yet every generated prose section type still read AI 81-100% of the time. The detector here was trained on that pipeline's output, so it recognises the pipeline's fingerprint by construction.
Why would a decision framework read AI every time? One plausible explanation is that a fixed shape leaves the model little to draw on beyond general knowledge, while a real author profile supplies facts about one person. Our data shows where the shapes concentrate; it does not prove that mechanism.
The clearest difference inside body sections came from first-hand records, on one site. For a sports equipment manufacturer whose articles are built on its own sales and shipping records, record-heavy sections read AI 37% of the time, against 71% for the other sections of the same articles. The limit: for a payments company whose articles used general industry figures, the split was 98% against 100%.
General-purpose detectors see much less of this. Four open detectors (desklib, Fakespot, Binoculars and Fast-DetectGPT) called 0 of 223 held-out articles from our pipeline AI at a 1% false-positive threshold, while catching 13-80% of ordinary public AI text. A classifier trained on our pipeline's output separated the same articles from human writing perfectly (AUROC 1.000), so the fingerprint is learnable. Commercial detectors such as GPTZero and Pangram were not part of that test.
Newsrooms have reached a related conclusion from the policy side. A Nieman Lab analysis by Hannes Cools and Nicholas Diakopoulos, covering 21 newsroom AI guidelines published by July 2023, found approaches ranging from outright bans (Wired does not publish AI-generated text, except when that is the point of the story, or text edited by AI) to conditional rules (The Guardian requires senior editor permission). The Dutch news agency ANP describes its process as "Human>Machine>Human." The authors note that "it is often far from clear from the guidelines how these mentions of transparency will take shape."
Guidelines like these govern who approves and labels AI output. Whether a section reads generic depends on what it was built from, and that leaves one question an editor can ask of every section: does it contain information only this organization could publish?
Forward Signal - 12-24 months horizon
Where AI-Text Detection Is Headed
Three forecasts on how detection methods, labeling policies, and humanization services evolve over the next 12-24 months.
Three Predictions For Detection and Disclosure
Use these predictions to judge which detection and disclosure trends are likely to affect writers and publishers next.
Over the next 12-24 months, detection methods will increasingly target structural patterns (sentence rhythm, section type and punctuation density) rather than banned vocabulary lists, since read-AI rates in one pipeline's articles already range from 8% in author bios to 100% in decision-framework sections, and one company's AI drafts carried about 16 times the dash punctuation of its own pre-AI blog (12.9 vs 0.8 per 1,000 words).
Rather than publishers adopting stricter word-ban style guides, demand will grow for paid manual humanization services and prompt-engineering workarounds, since labeling content as AI-written does not significantly change how persuasive readers find it and detectors remain easy to fool.
Institutions, from the U.S. Copyright Office to newsrooms and universities, will keep expanding AI-disclosure and authorship policies over the next 12-24 months even though the detectors meant to enforce them remain unreliable, following the pattern of Amazon's KDP disclosure requirement, Ringier's labeling rules across 19 countries, and one Salem State University writing instructor's near-total ban on generative AI in a course.
Faint signals worth tracking: In our scoring of one pipeline's articles, read-AI rates already vary from 8% to 100% by section type, and in Epoch AI's test of three commercial detectors, AI text written to mimic a specific author's style went undetected in about 13% of passages. The Copyright Office already evaluates AI authorship case-by-case, Amazon KDP requires disclosure of AI-generated (not AI-assisted) content, 21 newsroom AI guidelines have been catalogued across the U.S. and Europe, and one Salem State University writing instructor bans generative AI for almost all writing in a course.
Evidence Behind Each Prediction
Each forecast lists the real-world data points that support it alongside findings that complicate it.
- AI detectors rarely flag human writing, but sometimes miss AI text is the strongest public backing for this call. [Industry Publication]Epoch AI tested three AI text detectors (Pangram, GPTZero and Originality.ai) on both AI-generated and human-written text (Data Insight, Jul. 15, 2026). “The pattern across conditions is consistent: false-positive rates are low overall, with only Originality.ai mistaking any human-written text for AI.”
- The case rests on Prompt to make AI content not sound like AI content? [Community / Forum]Post originated in r/PromptEngineering, submitted by u/resiros approximately 1 year before the current context date (thread age listed as "1y ago"). “You are a writing assistant trained decades to write in a clear, natural, and honest tone.”
- Labeling AI-Generated Content May Not Change Its Persuasiveness supports this forecast. [Academic]Study surveyed more than 1,500 U.S. participants on perceptions of AI-generated policy messages. “There is good reason to assume that AI labels make people more skeptical of the underlying content.”
- The Problems with AI Detectors: False Positives and False Negatives is what puts this forecast on the board. [Academic]Turnitin has stated its AI checker has a less than 1% false positive rate. “I could pass any generative AI detector by simply engineering my prompts in such a way that it creates the fallibility or the lack of pattern in human language.”
- The case rests on 25 Tips to Humanise AI-written text and avoid AI detection. [Video]The speaker positions themselves as working directly with AI tools and helping develop some of them, in addition to editing AI-written text professionally. “Do not use AI humanizer tools if you don't have to.”
- AI Copyright: Lesson Learned - by John Overman - Daily Delirium is what puts this forecast on the board. [Substack / Newsletter]U.S. Copyright Office's "Copyright and Artificial Intelligence, Part 2: Copyrightability" (page iii) states copyright protects original human expression even if AI-generated material is included, but does not extend to purely AI-generated… “The use of AI tools to assist rather than stand in for human creativity does not affect the availability of copyright protection for the output.”
- Writing guidelines for the role of AI in your newsroom? Here are supports this forecast. [Industry Publication]Analysis covers 21 newsroom AI guidelines from organizations in the U.S., Europe, and elsewhere, compiled by Hannes Cools and Nicholas Diakopoulos, published July 11, 2023 on Nieman Lab. “all material published has been reviewed by a human and falls under our publishing authority.”
- Teaching Writing Tip of the Week: Banning AI (Almost Completely) in points the same way. [Substack / Newsletter]Salem State University Writing Center policy, authored by "Amy," published Sep 11, 2024, bans generative AI (e.g., ChatGPT) for "ALMOST ALL" writing in the course.
Conditions That Would Reverse These Trends
These scenarios describe the conditions that would push detection and disclosure trends in a different direction.
Before you rely on these numbers
A score measures how much current evidence backs a call, and that evidence keeps moving. The top forecast here sits at 95/100, while the minority view at 95/100 shows where the sources still disagree.
- Detection Moves From Vocabulary To Structure. Buyers changing priorities, or regulators changing rules, hit that call first.
- Humanization Services Outgrow Word-Banning Compliance. A source base that turns contrary would leave that as the forecast still standing.
In July 2026 we measured 816 AI-drafted articles across 24 websites, and explicit bans worked literally: banned words appeared in 0-12 of them (98-100% compliance). In a separate September test, with shared phrases down to 0.5% or less of each section, 81-100% of generated prose sections still read AI to a detector trained on that pipeline.
What do the "AI words to avoid" lists actually measure?
Shared phrasing across a fleet of articles. The lists catch wording that many AI drafts reuse, but they say little about the underlying model style that makes a passage read machine-written.
The 816-article corpus shows how visible fleet phrasing is. One stock opening phrase appeared in 74% of articles, on every site, and em dash characters appeared in 34%. You can count patterns like these and report a number, which is one plausible reason word lists became a default quality check for large pipelines. Model style cannot be counted the same way.
I think of this as the two-question test: one question about the fleet fingerprint, one about model style. A compliance audit that confirms the banned words are gone answers the first question. It does not answer the second.
| Question | What it measures | What phrase bans address |
|---|---|---|
| Fleet fingerprint: does this read like thousands of other articles from the same pipeline? | Shared phrase frequency across a corpus | Yes. Banned words appeared in 0-12 of 816 articles (98-100% compliance). |
| Model style: does this prose read machine-written to a trained detector? | Section-level readings from a detector trained on that pipeline | No. With shared phrases at 0.5% or less of each section, 81-100% of generated prose sections still read AI. |
Phrase lists answer the first question well. They say little about the second.
What reduced shared wording further was not a longer ban list. When each article was assigned its own sampled opening move, rhythm quota and phrasing, verbatim 8-word overlap with the rest of the corpus fell from 60 shared phrases in the control to 29, then to 18, and to 2 once per-site phrasing pools and a structure sampler were added. The channel mattered as well. The same style instruction was partly ignored when returned inside a tool response, followed almost perfectly when embedded in the writing skill, and followed best when delivered as per-section delivery notes.
Instructions can also create the tell they are meant to remove. A "humanizer technique" that mandated negative parallelism ("it's not X, it's Y") was itself one of the best-known AI tells. A rule that asked for dash asides next to a ban on em dashes produced spaced-hyphen asides instead.
Labels and wording also sit apart from how readers respond. In a Stanford HAI policy brief on a study by Isabel Gallegos, Robb Willer and colleagues, more than 1,500 U.S. participants read GPT-4o policy messages presented as written by an expert AI model, by a human policy expert, or with no authorship information. The label changed whether people thought the author was AI or human. It did not significantly change how persuasive they found the messages. The study covered four less-polarized policy topics and measured persuasion, not detection, so it supports a narrow point: a surface cue about origin did not change how readers weighed the argument.
The detector behind our section readings is a ModernBERT classifier plus a separate predictability check from a pair of open Qwen2.5 models. It scores passages of about 300 words, calls AI at 0.998 or higher and human below 0.5, and leaves everything in between not called. Its false-positive cap is at most 0.5% of human documents in every writing register, and on a 4,546-document pool of verified human writing it called none AI.
The clearest difference we found came from evidence, not wording. On one sports equipment manufacturer's site, article sections built on its own sales and shipping records read AI 37% of the time, against 71% for the other sections of the same articles. For a payments company whose articles used general industry figures, the same comparison came out at 98% against 100%. One site is not a rule, but it points at what the sections were built from rather than which words they used.
What will matter most for AI content compliance in the next 12 to 24 months?
We expect section-level style signals, disclosure rules tied to who created the content, and first-hand evidence to matter more than phrase bans. These are expectations, not measured outcomes.
The phrase-ban instinct treats detection as a vocabulary problem and disclosure as a detection problem. Neither holds up against the evidence above.
| Signal | What is already visible | Why it matters for your content team |
|---|---|---|
| Detection shifts from vocabulary to structure. We expect detection work to focus on section type, rhythm and style imitation rather than on word lists. | In our data, read-AI rates ranged from 8% to 100% by section type while shared phrases stayed at 0.5% or less of every section. In Epoch AI's July 2026 test of three commercial detectors, AI text from basic prompts was missed at most 0.7% of the time, but about 13% of passages written to imitate a specific author went undetected (38 of 297 on average), and about 26% in scientific writing. | Style guides built on phrase lists will not keep pace with style imitation or with detectors trained on a pipeline's output. Teams will need section-level checks, not just article-level phrase scans. |
| Disclosure rules expand independent of detection accuracy. We expect platforms and publishers to keep drawing the authorship line at creative origin, not at detector output. | Amazon KDP already requires disclosure of AI-generated text but not AI-assisted text. The U.S. Copyright Office analyzes human authorship case by case. Nieman Lab catalogued 21 newsroom AI guidelines by July 2023, with transparency and labeling as recurring themes. | A strategy built around lowering detector scores will not satisfy a platform that asks who created the draft. Teams need clear records of what was authored and what was generated and then edited. |
| Demand grows for substance over phrase compliance. We expect investment to shift toward content that carries first-hand evidence rather than synonym edits. | In the Stanford HAI study, AI labels did not significantly change how persuasive readers found policy messages. On one manufacturer's site, sections built on its own records read AI 37% of the time against 71% for the rest of the same articles. | The clearest change in our readings came from first-party evidence: figures and outcomes specific to one organization. That evidence does not change the disclosure duty for AI-drafted text. |
A common prediction is that detection will eventually catch everything and force universal disclosure. The evidence so far is mixed. Commercial detectors in Epoch AI's test caught plainly prompted AI text almost every time, yet style imitation still slipped past them, and the RAID benchmark (ACL 2024) found detectors easily fooled by adversarial attacks, sampling changes, repetition penalties and unseen models. We expect teams that solve the substance problem first to be less exposed whichever way detection goes.
Questions this article answers:
- Do AI words-to-avoid lists actually prevent AI detection?
- Why does my content still fail AI detection after removing all the flagged words?
- What actually makes writing read as AI-generated after all the phrase edits are done?
The phrase-ban instinct is reasonable: AI generators produce repetitive wording at scale, and that wording is measurable and removable. Our bans reached 98-100% compliance. But compliance stops at the fleet fingerprint. The model style underneath stayed, and 81-100% of generated prose sections still read AI to a detector trained on that pipeline.
Prompt-level fixes have limits of their own. One experienced user in r/PromptEngineering writes that as a conversation moves out of the context window, a chatbot "will fade back to the defaults." In our pipeline, the same style instruction was followed best when delivered as per-section notes rather than returned inside a tool response.
Disclosure is a separate line again. Amazon KDP counts text an AI tool created as AI-generated even after substantial edits, so neither phrase removal nor first-hand evidence changes that label for an AI-drafted book. What first-hand evidence changes is the substance: on the one site where sections used the company's own records, those sections read AI 37% of the time against 71% for the rest.
That is what the AEO Content Engine is built for: articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them.
This article is part of our research series on how AI writes and how humans write. The overview of the whole series is How AI Writes vs How Humans Write.
To see whether a draft still reads machine-written after the phrase edits, run it through the free AI Content Detector.
Written by
Alex Shortov
CTO, AEO Content
Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.
Connect on LinkedInSummarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently asked questions about AI words to avoid in writing
Teams ask these questions most often after applying phrase bans to their own drafts and checking what changed and what did not.
Does removing AI buzzwords from a draft prevent AI detection?
No. In our July 2026 measurements, bans reached 98-100% compliance across 816 articles. In a separate September test, with shared phrases at 0.5% or less of each section, 81-100% of generated prose sections still read AI to a detector trained on that pipeline.
What is the difference between AI-generated and AI-assisted content?
Under Amazon KDP's guidelines, content is AI-generated when an AI-based tool created it, even if you applied substantial edits afterwards, and it must be disclosed. AI-assisted content is created by the author and only edited, refined or error-checked with AI tools, and KDP does not require disclosing it. The line is creative origin, not vocabulary.
Do prompt instructions to avoid certain phrases work reliably over a long session?
Not reliably. In a r/PromptEngineering thread, one user writes: "You will never get to zero emdashes, and as you move out of the context window, your chatbot will fade back to the defaults." The same user suggests starting a new chat to reset. In our pipeline, instructions also executed literally: banning a heading moved the phrase into the first sentence.
Are commercial AI detectors reliable enough for content compliance decisions?
Not as the only gate. In Epoch AI's July 2026 test, Pangram and GPTZero flagged none of 495 human passages and Originality.ai flagged 19, but about 13% of passages written to imitate a specific author went undetected. The University of San Diego guide advises against using detectors as the sole indicator of academic misconduct.
What made AI-drafted sections read more human in your data?
First-hand records. On one sports equipment manufacturer's site, sections built on its own sales and shipping records read AI 37% of the time, against 71% for the other sections of the same articles. Where a payments company's articles used general industry figures, the split was 98% against 100%, so generic numbers did not help.
Why did banning a phrase make it show up somewhere else?
Because a ban executes exactly as written. When a stock opening phrase was banned as a heading, the model moved the phrase into the first sentence, and a generator built to vary wording later put it back until we added a phrase-level guard. Bans have to name the phrase itself, not just its position.