AEO Content AI: free AI content detector AEO Content AI: free AI content detector

Why Human Writing Gets Flagged as AI, and Which Writing Gets Flagged Most

An AI detection false positive is a human-written document that a detector calls AI. In our tests, the human writing misread most often was neutral institutional prose: a government program update, a policy analysis, a corporate case study.

Formal institutional policy documents on a desk beside a laptop showing text analysis, representing the register overlap between official prose and AI-generated writing that causes false positives in AI detection
Three things writers believe about AI detection. Myth or fact?
Call each one, then see how other readers called it.
1 AI detectors can reliably identify whether any text was written by a human.
2 Formal institutional prose gets flagged as AI more often than personal essays or tutorials.
3 Reducing false positives on human writing necessarily makes a detector miss real AI content.

Quick Answer

An AI detection false positive is a human-written document that a detector calls AI. In our tests, the human writing misread most often was neutral institutional prose: a government program update, a policy analysis, a corporate case study. One plausible explanation, not a finding, is that this is the register AI models imitate by default. An overall rate can hide the problem: our previous detector called 0.14% of pre-2021 Medium posts AI but 1.26% of US federal agency documents. Retraining with more agency prose brought the agency figure down to 2 of 637.

Did this answer your question?

Among human writing, the kind our AI detector misread most often was formal institutional prose: government updates, policy analyses and corporate case studies. Our previous detector called 8 of 637 US federal agency documents AI (1.26%), above our own 0.5% cap, while it called 4 of 2,868 pre-2021 Medium posts AI (0.14%). A plausible explanation is that neutral, well-organized explanatory writing is the register AI models imitate by default, but our data shows where the errors fell, not why.

Other writers point to editing tools. One writer on Medium described two drafts written without ChatGPT but edited with Grammarly that a publication flagged as AI. He could not tell which parts were flagged, and he presents Grammarly as a suspicion rather than a confirmed cause. We have not measured editing tools, so treat that as an open question.

Stated error rates and reported ones do not always match. A University of San Diego legal research guide, The Problems with AI Detectors: False Positives and False Negatives, notes that Turnitin has stated a false-positive rate of less than 1%, while a Washington Post study found 50% on a much smaller sample. The same guide calls AI detectors "problematic and not recommended as a sole indicator of academic misconduct."

Reported cases show what is at stake. In a Reddit thread titled Received a zero because my essay was flagged as AI-Generated, a student describes getting a zero on a self-analysis essay after Scribbr's AI detector flagged it as "39-100% AI-Generated," adding: "I am almost 100% certain that my professor didn't even read my essay." In that account, the score was treated as the evidence and the writer's word carried little weight.

Our own detector is tested against a cap: at most 0.5% of verified human documents may be called AI in every type of writing, not just on average. When agency prose broke that cap, we added 100 documents from other agencies to the training set. The retrained version called 2 of the 637 agency documents AI and still caught 99.6% of held-out AI articles. The sections below give the numbers, the fix and the limits.

Across 9,963 verified human documents, the current version of the AEO Content AI Content Detector called 4 AI. The clearest pattern came from its predecessor, and it was about which human writing gets misread. Most of the false positives were neutral institutional prose: three of the four pre-2021 Medium posts the previous detector called AI (a USAID program update, an Affordable Care Act policy analysis and a technology company's design case study), and 8 of 637 US federal agency documents.

We cannot prove why. One plausible explanation is that instruction-tuned models default to that register. Reinhart et al. (PNAS, February 2025) found model writing to be noun-heavy and information-dense, with a larger gap from human writing for instruction-tuned models than for base models. If explanatory institutional prose leans the same way, a detector has less to separate the two, and we call that possible overlap the register gap.

Formal writing is not the only group at risk. In a study reported by Stanford HAI, seven widely used detectors were near-perfect on essays by US-born eighth-graders but classified 61.22% of TOEFL essays by non-native English writers as AI. All seven detectors flagged 18 of the 91 TOEFL essays (19%), and 89 of them (97%) were flagged by at least one.

The false positives we found were fixable. After we added 100 agency documents to training, the retrained detector called 2 of the 637 agency documents AI and 0 of 511 documents from eight agencies it had never seen, while still catching 99.6% of held-out AI articles.

Writer reviewing formal documents beside a laptop showing AI text analysis, illustrating how institutional prose triggers false positive flags in AI content detection tools
Formal writing registers, used in government reports, academic papers and policy documents, may overlap with the patterns AI language models produce by default.

The next 12-24 months, scored

Where AI Writing-Detection Accuracy Is Headed

Three forecasts on how detection bias against non-native and formal writing evolves over the next two years.

29 sources analyzed8 community discussions3 blog posts3 academic sources1 industry publication
A

Forecasts for Writing-Detection Accuracy

Use these forecasts to gauge how much a single AI-detection score should weigh in academic, hiring, or publishing decisions.

88/100
High confidence 12-24 months

Detectors that score writing by how predictable its wording is will keep misclassifying non-native English and highly formal prose as machine-generated at multiples of the average rate, sustaining disputes in academic and hiring settings through the next two years.

83/100
Medium confidence 12-24 months

Rising documented cases of students and freelancers losing grades, income, or platform access over disputed AI flags will push schools, publishers, and marketplaces toward requiring corroborating evidence before penalizing writers, rather than relying on a single detector score.

Faint signals worth tracking: Seven widely used detectors flagged 61.22% of TOEFL essays from non-native English writers as AI-generated, versus near-perfect accuracy on U.S.-born eighth-graders' essays. Across 9,963 verified human documents, AEO Content's own detector made 4 false positives, none of them on a 4,546-document pool (95% upper bound 0.08%); on 2,868 pre-2021 Medium posts its previous version called 4 AI (0.14%) and the current version 2. Those detectors and test sets differ from the TOEFL study, so the rates are not directly comparable. A student received a zero grade after a detector flagged a self-analysis essay as 39-100% AI-generated, while Turnitin's own stated false-positive rate of under 1% diverged sharply from a Washington Post study that found 50% in a smaller sample.

B

Supporting and Contrary Evidence

Each forecast lists the real-world data points that support it and the findings that could challenge it.

Non-native and formal writing keeps drawing false AI flags 88
Supporting evidence
  • AI-Detectors Biased Against Non-Native English Writers | Stanford HAI is the strongest public backing for this call. [Academic]At least seven developers/companies have launched AI detectors following ChatGPT's high-profile launch. “These numbers pose serious questions about the objectivity of AI detectors and raise the potential that foreign-born students and workers might be unfairly…”
  • Guilty Until Proved Human - Slow AI supports this forecast. [Substack / Newsletter]Weber-Wulff et al. (2023) tested 14 AI detection tools; none scored above 80% accuracy. “AI detection tools do not work. They discriminate against non-native English speakers. And universities are still paying millions for them.”
  • The Problems with AI Detectors: False Positives and False Negatives is what puts this forecast on the board. [Academic]Turnitin previously stated its AI checker had a less than 1% false positive rate. “We're comfortable with that [false negative rate] since we do not want to highlight human-written text as AI text.”
False-flag disputes push institutions toward less punitive use of detection scores 83
Supporting evidence
  • The case rests on Received a zero because my essay was flagged as AI-Generated. [Community / Forum]Original poster (u/Tokkishin) received a zero on their Final Self-Analysis essay after Scribbr's AI-Detector flagged it as "39-100% AI-Generated.". “I am frustrated because I am almost 100% certain that my professor didn't even read my essay.”
  • The Problems with AI Detectors: False Positives and False Negatives is what puts this forecast on the board. [Academic]A Washington Post study found a much higher false positive rate of 50%, though with a much smaller sample size than Turnitin's.
  • Turnitin flagged my human written work as AI written supports this forecast. [Community / Forum]Original poster (OP) states Turnitin flagged 60% of their self-written meta-analysis as AI-generated after a month of independent work. “It's really a shame the educational instiutions are using what is basically high tech astrology to determine people's grades.”
C

What Could Change These Forecasts

These scenarios describe the conditions that would most likely shift how AI-detection tools are built or used.

A note on uncertainty

Treat these scores as weights, not verdicts. The top signal (95/100) leads on evidence, and the minority view (95/100) marks where sources spread out.

  • Calibration, not detection itself, is the fixable problem. Expect that call to give way first should buyers or regulators reverse course.
  • Calibration, not detection itself, is the fixable problem. Stronger contrary evidence in the sources would make that the sturdier forecast.
Methodology Each signal scored 0-100 by an evidence-weighted model based on source authority, recency, and how many sources support it.

What Will Determine Which AI Detectors Remain Credible Over the Next 12 to 24 Months?

We expect credibility to follow false-positive rates published for each kind of writing and checked on unseen text. Scores used alone against writers will keep drawing disputes. Both are forecasts, not findings.

Signal What the evidence shows now Why it matters over the next 12 to 24 months
Non-native and formal writing will keep drawing false AI flags from detectors not tested on it In a Stanford study by Liang et al. (2023), seven widely used detectors classified 61.22% of TOEFL essays by non-native English students as AI, while they were near-perfect on essays by US-born eighth-graders. In our own tests, our previous detector called 1.26% of US federal agency documents AI, above our 0.5% cap. Where a single detector score is grounds for a zero grade or a rejected submission, writers in these groups carry more of the risk. As disputes accumulate, the gap between a tool's headline false-positive rate and its rate on specific kinds of writing becomes harder to dismiss.
What a detector is trained and tested on decides where it errs Adding 100 documents from other agencies to our training set cut agency false positives from 8 to 2 of 637, with 0 of 511 on eight agencies never used anywhere, and the detector still caught 99.6% of held-out AI articles. Adding too much human text can backfire: 404 marketing posts once cost 13-27 points of recall. If testing by kind of writing becomes a standard expectation, a single headline rate will carry less weight as grounds for treating a score as conclusive. The question shifts from "was this written by AI?" to "was the detector tested on writing like this?"
Documented false-flag disputes will push institutions toward corroboration requirements A student on Reddit describes receiving a zero after Scribbr's AI detector flagged a self-analysis essay as "39-100% AI-Generated." A University of San Diego legal research guide sets Turnitin's stated false-positive rate of less than 1% against a Washington Post study that found 50% on a much smaller sample, and calls detectors "problematic and not recommended as a sole indicator of academic misconduct." Schools and publishers acting on a single detection score will face pressure to ask for corroborating evidence, such as drafts, edit history or archive dates, before imposing penalties, especially for non-native English writers and writers working in formal registers.

I would push back on the framing that treats all AI detection as fundamentally broken. That claim is too broad to be useful. In our tests the misreadings were concentrated rather than spread evenly: none in tutorials or personal essays, most in institutional prose, and a targeted training fix cut them on sources the model had never seen. That is a narrower claim than "detection works," and it comes with limits. These are internal results, not an independent audit; the language models used for scoring may have memorized old public text, which can make human writing look more predictable to them; and a set of never-published human writing, the strongest proof, is still to come.

Publishing formal writing that has to hold up?

The AEO Content Engine grounds every article in your company's own records and writes it to read like careful expert prose, structured for AI answer engines to cite.

Top Questions This Article Answers

These are the exact phrases writers, students, non-native English speakers, and content teams are searching to understand why human writing gets flagged, which registers produce the most false positives, and what register-specific calibration does to the problem.

Which Human Writing Gets Flagged as AI Most Often?

Neutral institutional prose. Of 2,868 pre-2021 Medium posts, our previous detector called 4 AI (0.14%); three were a USAID update, a policy analysis and a corporate case study.

The fourth was a list of 100 headline-style post titles, not running prose. There were no false positives in tech tutorials, data science, crypto and finance, startups, personal essays or culture writing, with 350 posts in each of the five largest groups. On 1,401 fresh Medium posts from authors outside the first sample, the previous detector called 3 AI and the retrained version 0. The current version calls 2 of the original 2,868.

The agency documents made the pattern sharper. We collected 637 explanatory documents archived before 2021 from US federal agencies the detector had never seen, including USAID, the State Department's ShareAmerica, CFPB, VA, the Department of Labor, the Census Bureau and USGS. The previous detector called 8 of them AI (1.26%), above our 0.5% cap. ShareAmerica, plain English written for international readers, was the hardest source: 5 of its 90 documents were called AI. Our data does not settle why such writing trips a detector. A plausible explanation is that public communications aim to be neutral, precise and accessible, which is close to the default voice of instruction-tuned models.

Formal register is not the only risk factor. Stanford HAI reported a study by Liang et al. (2023) in which seven widely used detectors classified 61.22% of TOEFL essays by non-native English students as AI, while they were near-perfect on essays by US-born eighth-graders. All seven detectors flagged 18 of the 91 essays (19%), and 89 (97%) were flagged by at least one. The researchers link the gap to how those detectors score predictability: non-native writers tend to use less varied vocabulary and simpler sentence structure. Those detectors and essays differ from our test sets, so the rates are not directly comparable with ours.

Surface habits are a separate question from register. In one company's AI-drafted articles we measured 12.9 dash-punctuation marks per 1,000 words, against 0.8 in the same company's own pre-AI blog, and an outside AI-detection review of that site picked up the pattern. After we set a budget of 2 dash marks per 1,000 words and rewrote asides as commas, parentheses or separate sentences, new articles measured 1.69. A habit like that marks the AI drafts, not the human writing, so it does not explain why careful human prose gets flagged.

Our approach keeps two readings distinct. Origin (who most likely wrote the text) and style (how generic or distinctive it reads) are separate readings that we never blend into one score. A piece can be human-written and generic, or AI-drafted and distinctive. Our old style score shows why the split matters: on our own content it was inverted, scoring AUROC 0.22 on human client archives versus AI pipeline output (0.5 is a coin flip), because it measured the same surface features that humanizing optimizes, so humanized AI text looked more human than real humans. A style score is not evidence of authorship.

In our data, formal institutional prose was the human writing most often misread, and personal essays, tutorials and culture writing were not misread at all. That is a pattern in two archives, not a law, but it is where the risk showed up.

Why Did Retraining on Institutional Prose Cut False Positives Without Losing AI Detection?

Adding 100 documents from other federal agencies to training cut false positives on 637 agency documents from 8 to 2, and the retrained detector still caught 99.6% of held-out AI articles.

The usual objection is that teaching a detector to accept formal human prose lets AI writing in the same register slip through. Here it did not: the 99.6% comes from held-out AI articles the model had not seen in training. The trade-off is real in other settings, though. When we once added 404 human marketing posts as training negatives, human false positives went to zero but recall on AI articles fell by 13-27 points, while 50 posts cost nothing. We now tune that amount with recall as the gate.

Because we trained and compared five candidate models, we confirmed the winner on data that took no part in choosing it. On 511 documents from 8 agencies never used anywhere (Army, Air Force, Department of Transportation, FEMA, NIST, Department of Defense, Fish and Wildlife Service, Bureau of Land Management), the old detector made 3 false positives and the new one 0. On 1,401 fresh Medium posts from new authors, the result was the same: 3 before, 0 after. The fix held on writing the model had not been tuned on.

Published error rates deserve the same scrutiny. The University of San Diego guide The Problems with AI Detectors: False Positives and False Negatives reports that Turnitin has stated a false-positive rate of less than 1%, while a Washington Post study found 50% on a much smaller sample. It also notes that Turnitin's checker can miss roughly 15% of AI-generated text in a document, and that the company says it is comfortable with that because it does not want to highlight human-written text as AI. The same guide says studies indicate that neurodivergent students and students writing in a second language are flagged at higher rates. A single published rate does not tell a writer what to expect for their own kind of writing.

Dates are a separate problem, and they cut the other way. In one payments company's blog, 7 of 50 posts from January to October 2022, before ChatGPT's release, already read like GPT-3 or Jasper output. For marketing copy, "before ChatGPT" is not "before AI", so we only trust human labels for text dated 2021 or earlier. Rebuilt websites add another trap: when a site is migrated, old posts may be rewritten while keeping their original dates.

That is why we treat proof of human authorship as its own task. We take human text only from archives that saved the page at the time, such as Wayback Machine and Common Crawl captures up to 31 December 2020, never from a live page whose date may have survived a rewrite. A detector score carries no information about when a document was written or under what conditions, so a decision that rests on the score alone rests on partial evidence.

The conclusion we draw is narrow. Where a detector misreads a type of human writing, more training examples of that writing, confirmed on sources the model never saw, can cut the error without giving up recall, as it did here. It is not guaranteed, it can cost recall if overdone, and it does not replace provenance evidence such as capture dates and drafts. None of that requires believing detection is useless. It requires treating a score as a measurement with stated conditions and documented limits, not a verdict.

What Does It Actually Mean When an AI Detector Flags Your Writing?

A flag means the text scored like AI writing to that detector, not that a machine wrote it. In our tests, the human writing misread most often was neutral institutional prose.

The numbers show how uneven that risk is. Our previous detector called 1.26% of US federal agency documents AI (8 of 637) and 0.14% of pre-2021 Medium posts (4 of 2,868), with no false positives at all in tech tutorials, data science, crypto and finance, startups, personal essays or culture writing. We call the distance between that formal, information-dense register and more personal writing the register gap, as of . It is a way to frame the pattern, not a measured cause.

Three situations show up in the sources we reviewed. The first is formal institutional prose, which is what our own data shows. The second is writing by non-native English speakers, documented in a Stanford study of TOEFL essays. The third, reported by writers but not measured by us, is text passed through an editing tool: in one Reddit thread, a commenter blamed Grammarly for swapping plain words for "flowery" synonyms. Each can produce human text that a detector misreads.

A detector does not assess your intent or your process. It scores the text that results. When that text resembles what the detector learned to treat as AI writing, it can be flagged regardless of who typed it. That is a property of the detection method, not a judgment of the writer.

Detectors also behave differently from one kind of writing to another. When Stanford researchers built the DetectGPT method, they found that OpenAI's earlier pre-trained detector worked well on English news articles, performed poorly on PubMed articles and failed completely on German-language news, as described in Human Writer or AI? Scholars Build a Detection Tool (Stanford HAI). A detector tested on one kind of text tells you little about how it will treat another.

The evidence inside a text may matter too, though our data here comes from AI-drafted articles rather than human writing. On one sports equipment manufacturer's site, sections built on the company's own sales and shipping records read AI 37% of the time, against 71% for the other sections of the same articles. That is a single-site result with a clear limit: for a payments company whose articles used general industry figures, the split was 98% against 100%. A plausible reading is that facts only one company could publish pull text away from generic phrasing, but the data does not prove it.

Our detector is a ModernBERT classifier plus a separate predictability check from a pair of open Qwen2.5 models, scoring passages of about 300 words. A document is called AI when a passage scores 0.998 or higher, human when no passage reaches 0.5, and otherwise it is not called. The rule we set is that at most 0.5% of verified human documents may be called AI in every type of writing, not just on average. On a pool of 4,546 verified human documents across 17 kinds of writing, the current version called none AI (95% upper bound 0.08%), and across all 9,963 verified human documents we tested it called 4. These are internal results, not an independent audit.

Origin and style are different readings. A well-edited AI article can read human. A formally written human article can read AI. Keeping those two dimensions separate is essential to understanding what a detection score actually measures.

The human writing that gets mistaken for AI is not random. In our tests it clustered in formal institutional prose, such as government updates, policy analyses and corporate case studies, while tutorials, personal essays and culture writing were not misread at all. For non-native English writers, the Stanford TOEFL study shows a separate risk with widely used detectors.

What I take from this work is more specific than the claim that detectors have a bias problem. Where our detector broke its own cap, on agency prose, a targeted training fix cut the error from 8 to 2 of 637 documents and to 0 on 511 documents from agencies it had never seen, while it still caught 99.6% of held-out AI articles. The fix was bounded and testable, and the results are internal, not an independent audit.

The harder problem is provenance. Dates do not prove human authorship: 7 of 50 posts from 2022 on one company's blog already read like GPT-3 or Jasper output. In our view, no threshold fixes that; capture dates, drafts and edit history do more.

Specific records may help from the other side too. On one manufacturer's site, AI-drafted sections built on its own sales and shipping records read AI 37% of the time against 71% for the rest, though on a payments company's site using general industry figures the split was 98% against 100%. That is what the AEO Content Engine is built around: articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them. To see how a draft of your own reads, run it through the free AI Content Detector.

This article is part of our research series on how AI writes and how humans write. The overview of the whole series is How AI Writes vs How Humans Write.

Written by

Alex Shortov

CTO, AEO Content

Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.

Connect on LinkedIn

Summarize This Article With AI

Open this article in your preferred AI engine for an instant summary.

Frequently Asked Questions

What is the register gap in AI detection?

The register gap is our name for the possible overlap between formal institutional prose and the default style of AI models. Reinhart et al. (PNAS, February 2025) found model writing noun-heavy and information-dense, with a larger gap from human writing for instruction-tuned models than for base models. A plausible explanation for our results is that explanatory institutional prose reads the same way. Our data shows the result rather than the cause: our previous detector called 1.26% of US federal agency documents AI against 0.14% of pre-2021 Medium posts.

Why is my formal writing flagged as AI even though I wrote it myself?

A flag on formal prose is weak evidence on its own. In our tests the human writing misread most often was neutral institutional prose, the kind found in government updates, policy analyses and corporate case studies. As guidance: ask which detector was used and what its false-positive rate is for your kind of writing, and keep what shows your process, such as drafts, notes, edit history or an archived copy with a capture date. I would not interpret a flag on formal prose as meaningful evidence of AI origin without that context.

Does Grammarly or editing software make writing more likely to be flagged?

It has not been shown. We have not measured editing tools, and our sources include no controlled test of them. One writer on Medium reported two Grammarly-edited drafts, written without ChatGPT, that a publication flagged as AI, and he framed Grammarly as a suspicion rather than a proven cause. A Reddit commenter blamed Grammarly for swapping plain words for "flowery" synonyms. Treat editing tools as a possible factor, not an established one.

Are AI detectors biased against non-native English speakers?

The best-known evidence says some widely used detectors are. In a Stanford study by Liang et al. (2023), seven detectors classified 61.22% of TOEFL essays by non-native English students as AI while being near-perfect on essays by US-born eighth-graders, and 89 of the 91 essays (97%) were flagged by at least one detector. A University of San Diego legal research guide also reports studies indicating that students writing in a second language are flagged at higher rates. A single detector score should not be treated as conclusive for these writers.

Read next

Researcher analyzing AI detection calibration data and false-positive rates across writing registers on dual monitors

AI Detector False Positives: What a 0.5% Cap Really Means

Overhead view of an editorial desk with manuscript pages, statistical charts on a laptop screen, and handwritten analysis notes

How AI Detectors Work and Where They Break

Researcher reviewing printed documents and notebooks alongside a laptop showing an archive timestamp interface

How to Prove a Text Was Written by a Human