How AI Writes vs How Humans Write: What Nearly 10,000 Verified Human Documents Showed Us
AI writing and human writing differ in origin, not in quality. Surface tells like dashes and stock phrases can be edited out. What stays is the model's underlying style and the absence of first-hand specifics.
On this page
Quick Answer
AI writing and human writing differ in origin, not in quality. Surface tells like dashes and stock phrases can be edited out. What stays is the model's underlying style and the absence of first-hand specifics.
In our research, AI-drafted articles with dash marks cut to 1.69 per 1,000 words still read AI in 81-100% of generated prose sections to a detector trained on that pipeline. Our detector called 4 of 9,963 verified human documents AI. Stanford researchers found readers identify AI text with 50-52% accuracy. Sections built on a company's own records read AI about half as often.
Can anyone tell whether AI wrote an article? We checked from the measurement side. Since late August 2026 we rebuilt the AI Content Detector behind aeocontent.com, tested it on 9,963 human documents we could date to before modern AI writing, scored 146 AI-drafted articles from 13 websites section by section, and ran four open detectors against 223 of our held-out articles.
This article reports what that work showed: how AI writing differs from human writing once the obvious tells are gone, why detectors and readers misfire, and what actually changes the result. Every finding comes with its sample size and its limits.
Why does telling AI-written from human-written content resist a reliable test?
Because the differences that last are not the ones people look for. Readers and checklists key on surface cues, and those cues are either shared by careful human writers or easy to edit away.
AI writing tools reached content teams before anyone agreed on what their output looks like. Checklists filled the gap: em dashes, stock words, long documents, perfect grammar. Each of those can be edited out in an afternoon, and several of them describe a careful human writer just as well.
Proving who wrote a text is harder than it looks, too. A publish date proves nothing, because websites rewrite old posts and keep the old dates. Our human reference set therefore comes from pages the Wayback Machine and Common Crawl captured up to 31 December 2020, plus client archives dated 2021 or earlier. We stop there on purpose: in one payments company's blog, 7 of 50 posts from January to October 2022 already read like GPT-3 or Jasper output.
That is why intuition was not our instrument. We track false-positive rates separately for each kind of writing, and we hold out whole sources: a website, an author group or a government agency is used either to train the detector or to test it, never both.
The next 12-24 months, scored
Where AI writing and detection head next
Three scored forecasts on how machine-written text, reader preference, and detection standards move over the next two years.
What shifts in AI versus human writing
Read each forecast as a near-term bet on where writing and detection move, weighted by the strength of its evidence.
The detection market moves toward thresholds set on large verified-human corpora with false-positive caps in every writing register, because generic tools misfire on neutral institutional prose. Expect a 0.5% false-positive ceiling across registers, tested on 9,963 verified human documents, to become the buying criterion rather than raw catch rate.
The most cited giveaways of machine drafting, heavy dash use and runaway length, lose reliability within 12-24 months as writers and tools adopt tighter style limits. Machine output measured at 7.2 to 12.9 dash marks per 1,000 words against near-zero in pre-AI human archives will converge as production practices cut new work below 2 dash marks per 1,000 words.
Resistance rooted in 'the machine writes worse' weakens over 12-24 months. Blind panels already misjudge authorship, with identification near a coin flip at 50-52% and 66% of one reader test picking the wrong source, while 76% of that panel preferred the machine-generated sample and separate readers preferred fine-tuned machine output on stylistic accuracy and quality.
Faint signals worth tracking: A deliberate dash budget cutting new articles from 12.9 to 1.69 dash-punctuation marks per 1,000 words, closing the gap with the 0.0 to 0.8 seen in human archives. In a moderated reader test, 76% preferred the machine-written sample and 66% of 21 participants misidentified which passage a human wrote. Retraining on 100 public-domain federal agency documents cut human-text false alarms from 8 to 2 of 637 agency documents and from 1 to 0 of 4,546 verified human documents while still catching 99.6% of held-out machine articles.
Sources behind each call
Each forecast lists both the studies that back it and the findings that cut against it.
- Backing it: What Happens When AI Writes Better Than Humans? [Blog]The author claims Claude currently writes better blog posts, Facebook ads, and emails than he does, and does it "10x faster.".
- Backing it: The Results Are In: The McLuhan Test and AI Versus Human Writing [Substack / Newsletter]66% of the 21 participants incorrectly identified the source of the writing samples. “AI carries a certain stigma - warranted to a degree, in my opinion - but that stigma interferes with our ability to evaluate its output objectively.”
- Backing it: Was this written by a human or AI? [Research]Study participants could distinguish human- vs. AI-generated text with only 50-52% accuracy, roughly a coin flip (Jeff Hancock, Stanford). “That's worrisome because it creates a risk that these machines can pose as more human than us.”
What could flip these calls
Scenarios in detector accuracy, model output, and reader trust that would reverse these forecasts.
A note on uncertainty
Treat these scores as weights, not verdicts. The strongest signal here scores 95/100, and the minority view (75/100) reflects a real spread in what the sources report.
- Calibrated detection becomes the standard. The moment regulators or buyers head the other way, that call is the exposed one.
- Reader preference tips to machine text. Should the evidence swing against the mainstream view, that forecast outlasts the rest.
What nearly 10,000 verified human documents actually show
Our detector called 4 of 9,963 verified human documents AI-generated (0.04%) in . With false alarms that rare, the real differences show up: not dashes or length, but origin.
On the largest pool, 4,546 documents spanning 17 kinds of writing, it called none (95% upper bound 0.08%). We then confirmed it on 1,912 documents from authors and agencies collected only after the model was chosen, and it called none of those AI either. That floor is what makes the rest of this article worth reading: a detector that misfires on people cannot tell you how AI writes.
We call our method the Two-Era Test: find a body of writing from the same organization that spans a clear before-AI and after-AI period, then compare what can be measured. In the two company blogs we ran it on, the visible differences were dash rate, document length and a burst of stock AI vocabulary. The origin reading kept separating the two eras after those visible differences faded.
What two eras of the same blog reveal
We reviewed the blog archive of a sports equipment manufacturer (industry only, no client named). One human author wrote 288 posts there between 2009 and 2018, and in 2026 the site added 30 AI-drafted buying guides. The 288 human posts averaged 0.0 em-dashes per 1,000 words and 492 words per post. The AI-drafted guides averaged 7.2 em-dashes per 1,000 words and 2,927 words. Our detector flagged 2 of the 288 human posts (0.7%) and 73% of the guides, yet both groups sat in a similar band on our style scale (57 against 68). Style overlapped. Origin did not.
A payments company's blog showed the same split over a longer run. Across 502 dated posts from August 2018 to September 2026, stock AI vocabulary rose from 4.0 words per 1,000 in the human era to 22.6 between October 2023 and April 2025, when 93% of posts were flagged. From August 2025 the AI posts became polished: the vocabulary fell back to 4.6 per 1,000, close to the human baseline, yet the detector flagged 100% of them. The same four author names appeared in every era, on human and AI posts alike.
Dashes behave the same way. AI-drafted articles on one company's site carried 12.9 dash-punctuation marks per 1,000 words against 0.8 in that company's pre-AI blog, even though a sanitizer had already removed every em-dash character. An outside review counted the spaced hyphens as the same tell, and it was right. When we replaced our writing rule with a budget of 2 dash marks per 1,000 words, new articles measured 1.69, much closer to the human baseline.
Dash rate and stock words make useful first screens and poor verdicts. An editor with a style guide can close either gap, and the text underneath still reads machine-written.
What our detector actually reads
Our detector is calibrated with a 0.5% false-positive cap in every writing register, not just on average. A detector checked only on pooled human text can meet its headline number while misfiring on one kind of writing. Ours did exactly that before we fixed it: the previous version called 8 of 637 US federal agency documents AI (1.26%), above the cap, while its rate on marketing blogs and client archives was 0.02%.
It does not count dashes or measure sentence length. The origin reading comes from a ModernBERT classifier trained to separate human and machine text, applied to passages of about 300 words, plus a separate predictability check from a pair of open Qwen2.5 language models. A document is called AI only when a passage scores 0.998 or higher, and human when no passage reaches 0.5. Everything in between is not called. Style, meaning how generic or distinctive the writing reads, is a second reading that we never blend into the first.
The common assumption is that AI gives itself away by writing worse. It does not. When we took 146 AI-drafted articles from 13 websites apart section by section, reused phrases made up 0.5% or less of every prose section, yet 81-100% of every type of generated prose section still read AI to our detector, which was trained on that pipeline's output. Fixed rhetorical shapes read the most machine-like of all: decision frameworks (100%), outlook charts (99%) and before-and-after comparisons (94%). The opening was the most machine-like part of an article, with 87% of first passages reading AI.
Editing out the tells changes how text looks, not where it came from. Quality and provenance are different axes.
| Signal | Human archive (pre-AI) | AI-drafted content | After editing or polish |
|---|---|---|---|
| Dash marks per 1,000 words (two company sites) | 0.0 to 0.8 | 7.2 to 12.9 | 1.69 after a dash budget |
| Average document length (words) | 492 | 2,927 | Not measured |
| Stock AI vocabulary per 1,000 words (payments blog) | 4.0 | 22.6 | 4.6, with 100% of posts still flagged |
Why do detectors and human readers both get the AI call wrong?
Both misjudge polished, neutral text. Our previous detector's rare errors on 2,868 pre-2021 Medium posts clustered in formal writing, and human readers in controlled tests do little better than chance.
That detector miscalled 4 of the 2,868 posts as AI-generated (0.14%), and 3 of the 4 were neutral institutional prose: a USAID program update, a policy analysis of an Affordable Care Act bill and a corporate design case study. The fourth was a list of 100 headline-style titles. We report that failure openly. The weak spot was one register, formal and well-organized writing, which is the register AI models imitate by default. Those were the posts that read most like what people imagine AI sounds like. The current version, retrained with 100 public-domain federal agency documents, called 2 of those 2,868 posts AI and none of 1,401 fresh posts from new authors.
According to Stanford HAI research by Jeff Hancock and colleagues, published in March 2023 (the paper is titled "Human Heuristics for AI-Generated Language Are Flawed"), participants judging online dating, professional and hospitality profiles could distinguish human from AI text with only 50-52% accuracy, roughly a coin flip. What made the finding notable was not the low accuracy itself. It was that participants leaned on the same cues: high grammatical correctness, first-person pronouns, references to family life and informal language were often attributed, incorrectly, to human writers.
Human intuition about authorship is not random. It is systematically wrong in the same direction.
Where open detectors pass and where they do not
We tested four open AI detectors that never saw our data (desklib, Fakespot, Binoculars and Fast-DetectGPT) against 223 held-out AEO Content articles, with each threshold set so that at most 1% of verified human documents would be called AI. None of the four flagged any of those articles, and at most 1 of 146 newer articles was flagged, while the same thresholds caught 13% to 80% of ordinary public AI text from the HC3, Beemo, MAGE and RAID datasets. The detectors scored our articles as more human than real human writing (AUROC 0.22 to 0.40, where 0.5 would mean they cannot tell the two apart). Two of the four are trained classifiers; Binoculars and Fast-DetectGPT are zero-shot methods that score how predictable the text is to a language model.
That is not the same as untraceable, and we do not claim it is. A classifier trained on our own pipeline's output separates those articles from human writing perfectly (AUROC 1.000), so a learnable fingerprint exists. General-purpose detectors simply do not carry it. Commercial detectors such as GPTZero and Pangram were not part of this test.
We publish what we measured, including where our own detector gets it wrong. A single headline false-positive rate can hide a register where a detector misfires, as agency prose showed for ours, so we report the rate for each kind of writing we tested.
No single detector number tells the full story. What matters is whether the tool was checked on text that matches the register you are screening. At the threshold that kept its overall human false positives to 1%, the open desklib detector called 4.1% of US federal agency documents AI.
The heuristics people rely on, and why they drift
The folk heuristics for spotting AI writing circulate widely: look for em-dashes, long documents, transition phrases like "moreover" and "furthermore," and perfectly grammatical prose. Some of those signals have real support at the population level. The problem is that style guides, deliberate editing and phrase-ban lists can close each gap in sequence, and the signals that remain after that kind of systematic editing are the deeper ones that require actual measurement tools.
What the data showed us from across our corpus is that origin and style are two different readings. A formally written human document can look machine-like on the surface without being machine-generated. A carefully prepared AI article built on a company's own records can read human in the sections that carry those records. Conflating these two axes is where both detection tools and human readers tend to go wrong.
Surface heuristics narrow the search. They do not close it.
What actually makes AI-assisted content read as genuinely human?
First-hand data that only one company could publish. On one manufacturer's site, sections built on its own sales records read AI 37% of the time, against 71% for the rest.
Those records were the company's own sales and shipping data, such as orders per state and models shipped per county. Same articles, same author profile, same pipeline. The sections built on records no model could have seen in training read AI about half as often. That is the most actionable finding from our section-level scoring, and it points away from style editing.
The finding has a limit. Numbers alone do not do it: for a payments company whose articles used general industry figures, the same comparison came out at 98% against 100%. The lever is first-hand data given to the model before it writes, such as outcomes from the company's own work. Synonym swaps, phrase bans and sentence-length tricks change none of that.
Why origin and style are two separate readings
We keep origin and style as two readings that are never blended in our evaluation work. A piece can read generically and still be human-written. A piece can be AI-drafted and still read human in the sections built on first-hand records. Treating those as a single axis is the core mistake in most discussions about AI writing quality.
Section types tell the same story. Across the 146 articles, author bios built from a real profile read human 51% of the time, calls to action 33% and curated resource lists 21%, against 7% for body sections and 3% for openings. What reads human is what the model did not write from scratch: a real person's background, a company's own offer, sources a person chose.
This is why first-party data is more than a citation strategy. It changes what the text is made of, not only how it is sourced. Original data is not a quality marker layered on top of AI output. In our scoring, it is the ingredient that moved the origin reading most.
What the research supports and what it does not
Kobak et al., in Science Advances in July 2025, analyzed more than 15 million PubMed abstracts and found a jump in "style words" after ChatGPT's release, estimating that at least 13.5% of 2024 abstracts were processed with language models. Reinhart et al., in PNAS in February 2025, found that model text differs from human text in grammar and rhetoric, with a noun-heavy, information-dense style, and that the gap is larger for instruction-tuned models than for base models. Both confirm the signal is real at population scale. Kobak's method estimates how common model-processed text is across a corpus, not who wrote a single document, a distinction that matters for anyone using detection tools operationally.
The RAID benchmark (ACL 2024) found that current detectors are easily fooled by adversarial attacks, changes in sampling strategy, repetition penalties and unseen models. The DIPPER paraphraser (NeurIPS 2023) cut DetectGPT's detection rate from 70.3% to 4.6% at a 1% false-positive rate, and sentence-by-sentence paraphrase remains the hard case for our detector too: it catches 31% of DIPPER-rewritten GPT-4 text. Attribution is a different matter. Sun et al. (ICML 2025) told apart text from ChatGPT, Claude, Grok, Gemini and DeepSeek with 97.1% accuracy, and the differences persisted after another model rewrote, translated or summarized the text.
What this external research collectively supports is a view I hold from our internal work: surface patterns are easy to disrupt, a model's underlying style is not, and a detector trained on a specific model or pipeline keeps finding it. The only lasting way to make AI-assisted content read as human is to give it human substance to work with: first-hand records, real experience and facts only you could publish.
Where does the evidence leave content teams and editors today?
The markers that last are not the surface ones. Measuring them requires verified human writing to calibrate against. The heuristics most teams rely on to screen AI-assisted content do not hold.
Human judgment sits near chance. Automated detection built on uncalibrated thresholds misfires on formal agency and policy writing; our own previous version called 8 of 637 US federal agency documents AI. And a threshold is never set once: each new model generation, and each new kind of human writing, can move it.
That is the frame I would apply going forward. The question worth asking is not "did a machine write this?" but "does this article contain information only this organization could publish?" First-party data at the section level is what moved detection scores in our internal evaluations, and it is also the part of an article that AI answer engines cannot find anywhere else.
The teams that distinguish themselves over the next few years will not be those that best conceal machine assistance. They will be those with the richest first-party evidence base: the proprietary data, the expert views, the case specifics. They will use AI to structure and scale that material, not to substitute for it.
That is the model the AEO Content Engine is built on: articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them.
Each finding above has its own article in this research series:
- Why Human Writing Gets Flagged as AI, and Which Writing Gets Flagged Most
- Do Humanized AI Articles Pass AI Detectors? We Tested Four
- Which Parts of an AI-Written Article Give It Away
- The Em Dash Is Not the AI Tell. The Dash Rate Is.
- Why Banning AI Buzzwords Does Not Make AI Text Read Human
- What Makes AI-Drafted Content Read Human: First-Hand Data
- How One Company Blog Went From Human Writers to AI, Year by Year
- How to Prove a Text Was Written by a Human
- How AI Detectors Work and Where They Break
- AI Detector False Positives: What a 0.5% Cap Really Means
To see how a draft of your own reads, run it through the free AI Content Detector, which gives separate origin and style readings for each passage of about 300 words.
Written by
Alex Shortov
CTO, AEO Content
Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.
Connect on LinkedInSummarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently asked questions
The questions we receive most often about AI and human writing cluster around three themes: detection reliability, reader perception, and what first-party data actually changes in AI-assisted content.
Can readers tell if something was written by AI?
Not reliably. In a 2023 Stanford study, people judging dating, professional and hospitality profiles told human from AI text with 50-52% accuracy, and in a small blind test 66% of 21 readers picked the wrong author. Readers lean on cues such as grammatical correctness and first-person pronouns, which do not track authorship. Editorial screening that relies on reading alone runs into that ceiling.
What patterns distinguish AI-drafted text from human writing?
The visible ones are dash rate, document length and stock vocabulary, and all three can be edited away. What remains is structural: fixed rhetorical shapes such as decision frameworks and outlook charts, openings that read the most machine-like, and a noun-heavy, information-dense style that research links to instruction-tuned models. A detector trained on a specific model or pipeline still picks that up after the surface tells are gone.
Do AI content detectors produce accurate results?
Only within the kinds of writing they were checked on. A detector can hold false positives to 1% overall and still misfire on one register: at that setting, the open desklib detector called 4.1% of US federal agency documents AI. Look for false-positive rates reported per register, confirmed on data that took no part in setting the threshold.
Does adding proprietary data change how AI-assisted content reads to a detector?
Yes, when the data is first-hand. On one manufacturer's site, sections built on its own sales and shipping records read AI 37% of the time against 71% for the rest. Where articles relied on general industry figures, the gap nearly disappeared (98% against 100%). The records have to be something only that company could publish.
Why do content teams keep misjudging AI authorship?
Because the cues people rely on, such as formal tone, polished phrasing and consistent structure, appear in expert human writing as readily as in machine output. The human writing most often mistaken for AI in our tests was neutral institutional prose: government updates, policy summaries and corporate case studies. Bylines do not settle it either: one company blog kept the same four author names on human and AI posts from 2018 to 2026.