How One Company Blog Went From Human Writers to AI, Year by Year
We traced how one company blog moved from human writers to AI, using 502 dated posts that a payments company published between August 2018 and September 2026, and compared it with a sports equipment manufacturer's site.
On this page
Quick Answer
We traced how one company blog moved from human writers to AI, using 502 dated posts that a payments company published between August 2018 and September 2026, and compared it with a sports equipment manufacturer's site. Each post got separate readings for origin detection, which estimates whether a model or a person most likely wrote the text, and for style, including how often typical AI vocabulary appears. On the payments blog, the share of posts flagged as AI went from 0% to 93% in under two years. Later, AI-vocabulary fell back close to the human era's level while every post in the final era was still flagged. The origin reading comes from the AEO Content detector, which reports origin and style separately and never blends them.
One payments company's blog kept the same four author names from 2018 to 2026. Behind those bylines, the share of posts our detector flagged as AI went from 0% to 93%, dipped briefly, then reached 100%.
This piece traces that change era by era, using the five periods in the data: human copywriting (August 2018 to March 2022), AI creeping in (April 2022 to September 2023), low-quality AI content (October 2023 to April 2025), a brief return to human writing (May to July 2025) and polished AI content (August 2025 to mid-July 2026). Each era has its own combination of post volume, AI-vocabulary and flagged share.
The timeline does not start with ChatGPT. MarTech's March 2023 review of five AI writing assistants notes that ChatGPT launched in November 2022 and sparked debate among marketers about AI replacing human writers, but the same review reports that Jasper launched in February 2021. On the payments blog, AI was already creeping in from April 2022. In a separate check of one payments company's blog, 7 of 50 posts from January to October 2022 already read like GPT-3 or Jasper output.
The data also separates two things teams often treat as one. AI-vocabulary rose to 22.6 per 1,000 words in the low-quality era and fell to 4.6 in the polished era, close to the human era's 4.0. The origin reading did not fall with it: 100% of polished-era posts were flagged.
Company blogs are not the only place this shift shows up. Kobak et al. (Science Advances, July 2025) analysed more than 15 million PubMed abstracts and estimated that at least 13.5% of 2024 abstracts were processed with LLMs. That is a population estimate, not a verdict on any single paper, but it shows how quickly AI-processed text became ordinary.
A second site gives a cleaner comparison of two authorships. On a sports equipment manufacturer's site, 288 posts written by one human author from 2009 to 2018 were flagged 0.7% of the time, while 30 AI-drafted buying guides from 2026 were flagged 73% of the time. Their style scores sat much closer together than their origin readings did.
The article also covers why bylines and word lists miss the change, how origin and style readings differ, and what could change over the next 12 to 24 months as fine-tuning gets cheaper and platforms set disclosure rules.
The question behind this analysis was not whether a company used AI on its blog, but whether the switch could be dated post by post from the text itself, without relying on bylines, publish dates or keyword lists.
Dates need care. A publish date is not proof of human writing: rebuilt websites rewrite old posts but keep the old dates, and AI writing tools were in marketing use well before ChatGPT's release in November 2022. When we build human test sets for our detector, we take text only from archive captures made up to the end of 2020 and stop human labels at 2021, never trusting the live page.
The origin reading comes from our AI detector: a ModernBERT classifier plus a separate predictability check from a pair of open Qwen2.5 models, run on passages of about 300 words. A document is called AI when a passage scores 0.998 or higher and human when no passage reaches 0.5; anything in between is not called. Style is reported as a separate reading and never blended into the origin call.
Two companies, two timelines, and one pattern worth testing on any blog: the look of the text and the origin reading can move apart. Bylines record who is credited, not how a post was produced.
Questions this article answers
- How much of a company blog can shift to AI-generated content within two years?
- Why do AI buzzword lists and bylines fail to detect AI-written posts?
- What is the difference between style detection and origin detection in AI content?
Why do bylines and vocabulary word lists fail to identify AI-written blog posts?
A byline records who is credited, not how a post was produced, and AI-vocabulary can fall back to human levels while posts still read AI to an origin detector.
Bylines were the first signal we checked on the payments company's blog. The same four author names appeared in every era from 2018 to 2026, on human-written and AI-written posts alike. Nothing in the bylines marked the shift in April 2022, the low-quality AI era or the brief return to human writing in mid-2025. A byline is an attribution field, not a production log.
Vocabulary lists catch loud AI and miss quiet AI. In the low-quality era, from October 2023 to April 2025, AI-vocabulary reached 22.6 words per 1,000, more than five times the human era's 4.0, and 93% of posts were flagged. That era was loud enough for a word list. In the polished era from August 2025, AI-vocabulary fell to 4.6 per 1,000, yet 100% of posts were flagged. Our research reads it plainly: word lists stop working once content is edited, and origin detection does not.
Punctuation tells behave the same way. In September 2026 we measured a fleet of AI-drafted articles that contained zero em-dash characters, because a sanitizer had replaced each one with a spaced hyphen, yet carried 12.9 dash-punctuation marks per 1,000 words, against 0.8 per 1,000 words in one company's own pre-AI blog. Swapping the character kept the habit. After we replaced the rule with a budget of 2 dash marks per 1,000 words and rewrote asides as commas, parentheses or separate sentences, new articles measured 1.69 dash marks per 1,000 words. A budget can fix a visible habit. On its own, it says nothing about who wrote the text.
Human writing gets misread too, and the cases are instructive. On 2,868 Medium posts archived before the end of 2020, the previous version of our detector called 4 of them AI (0.14%). Three were institutional prose (a government programme update, a health policy analysis and a corporate design case study), and one was a list of 100 headline-style titles. The current version calls 2. Our research reads this as a matter of register: neutral, well-organised institutional prose is the register AI models imitate by default. That is a plausible explanation, not something four posts prove.
Dates mislead in the same way bylines do. On one payments company's blog, 7 of 50 posts from January to October 2022 already read like GPT-3 or Jasper output, before ChatGPT's release. And when a site is rebuilt, posts dated years earlier may carry new text under the old date.
Anyone auditing a company blog therefore needs more than the most accessible signals. Bylines and vocabulary counts showed the loud era clearly and the polished era barely at all. Neither substitutes for an origin reading once the vocabulary has normalized.
What is the difference between style detection and origin detection, and which one holds up?
Origin estimates whether a model or a person most likely wrote a text; style measures how generic or distinctive it reads. On the sites we studied, origin separated authorships where style overlapped.
Style can change after drafting. Editors can replace repeated phrases and cut stock vocabulary, and newer models may write with fewer obvious tells. On the payments blog, AI-vocabulary fell from 22.6 per 1,000 words in the low-quality era to 4.6 in the polished era. The data does not show why; better models, heavier editing or both are plausible explanations. The flagged share went the other way, from 93% to 100%.
The clearest comparison comes from a sports equipment manufacturer's site, where one human author's archive and a set of AI-drafted buying guides sit side by side.
| Measure | Human posts, one author (2009-2018) | AI-drafted buying guides (2026) |
|---|---|---|
| Posts | 288 | 30 |
| Flagged AI by our detector | 0.7% (2 of 288), 0-2% in every year | 73% (September 2026 detector version) |
| Score on our style scale | 57 | 68 |
| Average length | 492 words | 2,927 words |
| Em-dashes per 1,000 words | 0.0 | 7.2 |
Both groups scored in a similar style band, 57 against 68, while the origin readings sat at 0.7% and 73%. Style overlaps; origin does not. That is why our detector shows origin and style as two separate readings instead of blending them into one score.
An origin reading is only useful if it rarely misreads people. We set a hard cap from the start: at most 0.5% of human documents may be called AI in every type of writing, not just on average. On 4,546 verified human documents across 17 registers, including client archives dated 2021 or earlier, archived marketing and SEO blogs up to 2020, open-web text and sports and health copy, our September 2026 detector called 0 AI (95% upper bound 0.08%). Across all 9,963 verified human documents we test on, it called 4.
The per-register cap exists because averages hide problems. The previous version of our detector called 8 of 637 US federal agency documents AI, 1.26% and above the cap, even though it stayed under the cap on Medium posts. After retraining with 100 documents from other agencies, the current version calls 2 of the 637.
General-purpose detectors are a different story. We ran 223 held-out articles from our own pipeline, which writes and humanizes long-form articles, through four open detectors: desklib, Fakespot, Binoculars and Fast-DetectGPT. With thresholds set so that at most 1% (and separately 0.5%) of verified human documents would be called AI, and that rate confirmed on 8,540 other human documents, they called 0 of 223 AI while catching 13-80% of ordinary public AI text. They ranked our articles as more human than real human writing (AUROC 0.22-0.40).
That does not make those articles untraceable. A classifier trained on our pipeline's output separates the same articles perfectly (AUROC 1.000), and commercial detectors such as GPTZero and Pangram were not part of the test. Outside research shows both sides. Sun et al. (ICML 2025) told ChatGPT, Claude, Grok, Gemini and DeepSeek apart with 97.1% accuracy, and the result persisted after rewriting, translation and summarization. Yet paraphrasing can still defeat detectors not built for it: in the DIPPER study (NeurIPS 2023), DetectGPT's detection rate fell from 70.3% to 4.6% at a 1% false-positive rate.
Material matters as well. In a separate September 2026 study of article sections, a sports equipment manufacturer's sections built on its own sales and shipping records read AI 37% of the time, against 71% for the rest of the same articles. For a payments company whose articles used general industry figures, the split was 98% versus 100%. That is one strong site result with a clear limit, not a general rule.
On the payments blog, origin held up where vocabulary did not: AI-vocabulary returned close to the human era while the flagged share stayed at 100%. Neither reading is complete alone, which is why we report both.
What will change most about AI content detection in the next 12 to 24 months?
We expect three pressures to shape origin detection: cheaper fine-tuning, platform disclosure rules and disputes over paying for training content. These are forecasts, not findings.
Our read is that the business pressure around detection may change faster than the technical picture. A team that keeps a 2024 detection setup unchanged may find it answering yesterday's questions. Here is how we read the three signals.
| Signal | Prediction (12-24 months) | Weak signal today | Why it matters for content teams |
|---|---|---|---|
| Cheaper model fine-tuning | Domain-specific model variants will multiply, and origin classifiers trained on general chatbot output will need retraining to keep up with them. | According to a Hugging Face post on async GRPO training, a rank-1 LoRA adapter for a 1.5B model is a few megabytes, against about 3 GB for the full model. The RAID benchmark (ACL 2024) found detectors easily fooled by models they had not seen. | Detection tools sold on precision against today's public models will need more frequent retraining, or they may miss the next wave of fine-tuned output. |
| Editorial disclosure pressure | More publishing platforms will set disclosure rules for AI assistance, making origin checks part of the path to publication rather than an optional audit. | In July 2023, Medium's policy update required AI-assisted stories to carry a disclosure within the first two paragraphs and stopped distributing fully AI-generated stories beyond the writer's own network. | If platforms require disclosure, teams will need a way to check what they publish. Teams that start now will face an easier transition. |
| Training data licensing | Disputes over paying for training content will push some publishers to track which individual posts were written by people. | In an r/Blogging thread, bloggers argue that AI companies should pay for site content, and one commenter points to Cloudflare's Pay Per Crawl feature, then in private beta. | If licensing terms come to depend on human authorship, the audit need moves from the blog level to the individual post. |
Precision against current model output erodes with each model update, so it is worth asking any detection vendor how often they retrain and what false-positive rate they hold on human writing. Our own answer to the second question is above: 4 of 9,963 verified human documents called AI.
How did one payments company blog shift from human writers to AI between 2018 and 2026?
We scored 502 dated posts from one payments company blog, August 2018 to September 2026. Flagged posts went from 0% in the human era to 93% at the low-quality AI peak, then 100%.
The shift did not move in one direction. The posts fall into five eras, and each has its own combination of post volume, AI-vocabulary and origin reading. Looking at the eras separately says more than any single figure for the whole blog. The table ends in mid-July 2026, so the most recent posts are not assigned to an era.
| Era | Date range | Posts | AI-vocab per 1,000 words | Share flagged AI |
|---|---|---|---|---|
| 1: Human copywriting | Aug 2018-Mar 2022 | 54 | 4.0 | 0% |
| 2: AI creeps in | Apr 2022-Sep 2023 | 79 | 5.4 | 30% |
| 3: Low-quality AI content | Oct 2023-Apr 2025 | 266 | 22.6 | 93% |
| 4: Brief return to human writing | May-Jul 2025 | 18 | 8.1 | 17% |
| 5: Polished AI content | Aug 2025-mid-Jul 2026 | 51 | 4.6 | 100% |
Three details stand out. Polished-era AI-vocabulary (4.6 per 1,000 words) sat close to the human era's 4.0. The polished era's flagged share (100%) was higher than the low-quality era's 93%. And the only fall in flagged share after AI arrived came when human writing briefly returned.
Era 1, human copywriting (August 2018 to March 2022). 54 posts, AI-vocabulary at 4.0 per 1,000 words, and no posts flagged. Every later era is measured against this baseline.
Era 2, AI creeps in (April 2022 to September 2023). 79 posts in a year and a half. AI-vocabulary barely moved, to 5.4, while the flagged share jumped to 30%. The era began months before ChatGPT's release in November 2022, which fits our separate finding that, on one payments company's blog, some posts from January to October 2022 already read like GPT-3 or Jasper output.
Era 3, low-quality AI content (October 2023 to April 2025). 266 posts, more than the first two eras combined, with AI-vocabulary at 22.6 per 1,000 words and 93% flagged. Low-quality AI is loud, and this is the era a vocabulary check handles well.
Era 4, a brief return to human writing (May to July 2025). 18 posts, AI-vocabulary down to 8.1 and only 17% flagged. For three months between two AI eras, most posts were not flagged. The data does not tell us why the blog changed course.
Era 5, polished AI content (August 2025 to mid-July 2026). 51 posts, AI-vocabulary back down to 4.6 and 100% flagged. Polished AI is quiet: close to the human era on vocabulary, and flagged on every post.
Practitioners name the risks they can see. Writing in MarTech in August 2026 about building an AI content pipeline, Tania Brown calls "robotic-sounding language and hallucinated 'facts'" the two main issues. Both show up in review. Neither tells a reviewer, once the vocabulary is polished, whether a model wrote the post.
Platforms have said as much. In its July 2023 policy update, Medium is for human storytelling, not AI-generated writing, Medium called AI detection tools "currently unreliable to the point of being unusable" and said its human editors and curators "often spot it instantly." Our blog data does not test that claim about editors. It shows what a measured origin reading adds when the vocabulary looks normal.
Controlled tests are less kind to unaided judgment. In a Stanford HAI study (Hancock et al., March 2023), people judging dating, professional and hospitality profiles told human from AI text with 50-52% accuracy, and wrongly read grammatical correctness, first-person pronouns, family references and informal language as signs of a human writer.
One finding cuts across all five eras. The same four author names appeared in every era from 2018 to 2026, on human-written and AI-written posts alike. The bylines did not change when the writing did, so they carry no diagnostic weight. What does carry weight is the combination of post counts, vocabulary and origin readings above.
The finding we keep returning to is that vocabulary and origin can come apart. On the payments blog, AI-vocabulary fell from 22.6 to 4.6 per 1,000 words between the low-quality and polished eras, while the flagged share rose from 93% to 100%.
Word lists and buzzword filters measure something real, and they would have caught the loud era. They would not have caught the quiet one. On the sports equipment manufacturer's site, style scores told a similar story: 57 for the human archive and 68 for the AI-drafted guides, against flagged shares of 0.7% and 73%.
These are two sites, not a survey, and detector readings carry error rates of their own, which we publish. The practical question for any content review process is which reading each tool actually gives, and whether that tool has been tested on human writing like yours.
A practical checklist for auditing a company blog (guidance, not findings):
- Split the archive into periods and compare them, rather than relying on one figure for the whole blog.
- Do not treat bylines as evidence of human authorship; on the blog we studied they stayed the same across every era.
- Do not trust publish dates alone. Rebuilt sites keep old dates on new text, and archived snapshots are better proof.
- Treat posts from 2022 onward as possibly AI-assisted, even those published before ChatGPT's release.
- Read origin and style separately. A vocabulary count that looks human says little about origin.
- Ask any detector what share of human writing like yours it calls AI before acting on its verdicts.
For new content, the AEO Content Engine produces articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them. To check existing posts, run them through the free AEO Content AI Detector.
This article is part of our research series on how AI writes and how humans write. The overview of the whole series is How AI Writes vs How Humans Write.
Written by
Alex Shortov
CTO, AEO Content
Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.
Connect on LinkedInSummarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently Asked Questions
What do content teams most often get wrong about AI detection on company blogs?
Many teams trust the visible signals, such as bylines and AI-sounding words. On the payments blog we studied, both missed the polished AI era that the origin reading flagged on every post.
How can I tell if a company blog has switched to AI-generated posts?
Compare periods, not single posts, and read origin and style separately. On the payments blog, AI-vocabulary gave a clear signal in the low-quality era (22.6 per 1,000 words) but looked close to human in the polished era (4.6), when every post was still flagged by the origin classifier. Treat publish dates with care, since rebuilt sites keep old dates on new text.
Do AI detectors still work on heavily edited AI drafts?
It depends on the detector. Four open detectors called none of 223 humanized articles from our pipeline AI, but a classifier trained on that pipeline's output separated them perfectly (AUROC 1.000), and commercial detectors such as GPTZero and Pangram were not tested. Paraphrasing remains the hard case: our own detector called GPT-4 text rewritten sentence by sentence with the DIPPER paraphraser AI only 18% to 31% of the time, depending on the version.
Can a blog keep the same author bylines while switching from human to AI content?
Yes. The same four author names appeared in every era of the payments blog from 2018 to 2026, on human-written and AI-written posts alike. Author bylines are attribution fields, not production logs. They do not update when the way content is produced changes.
Why would polished AI content score lower on vocabulary-AI signals than raw early AI drafts?
The data shows that it did, not why. AI-vocabulary was 22.6 per 1,000 words in the low-quality era and 4.6 in the polished era, close to the human era's 4.0. Better models, heavier editing or both are plausible explanations. The origin reading did not follow the vocabulary density down: 100% of polished-era posts were flagged.
Did the human writers ever come back?
Briefly. From May to July 2025, between two AI eras, the payments blog published 18 posts with AI-vocabulary at 8.1 per 1,000 words and only 17% flagged. From August 2025 the polished AI era began, and every post was flagged.
Can a publish date prove a blog post was written by a human?
No. Websites rewrite old posts and keep the old dates, and AI writing tools were in marketing use before ChatGPT's release in November 2022: on one payments company's blog, 7 of 50 posts from January to October 2022 already read like GPT-3 or Jasper output. For human test sets, we use archive captures made up to the end of 2020 and stop human labels at 2021.
Is a 100% flagged share proof that every post in an era was AI-written?
No single reading is proof. The detector scores passages of about 300 words and calls a document AI only when a passage scores 0.998 or higher, and any detector misreads some human writing. Ours called 4 of 9,963 verified human documents AI. A flagged share across 51 posts is strong evidence about an era, not a certificate for each post.
How often does the detector wrongly call human writing AI?
Rarely. Across 9,963 verified human documents, our detector called 4 AI. On a pool of 4,546 human documents across 17 types of writing, it called none (95% upper bound 0.08%). The human writing most often mistaken for AI is neutral institutional prose, such as government updates and policy summaries.