AI Detector False Positives: What a 0.5% Cap Really Means
An AI detector false-positive rate is the share of verified human documents a detector incorrectly calls AI. That rate can vary by type of writing: our previous detector called 0.14% of pre-2021 Medium posts AI but 1.26% of US federal agency documents.
On this page
Quick Answer
An AI detector false-positive rate is the share of verified human documents a detector incorrectly calls AI. That rate can vary by type of writing: our previous detector called 0.14% of pre-2021 Medium posts AI but 1.26% of US federal agency documents. The AEO Content AI Content Detector, a ModernBERT classifier plus a predictability check from a pair of open Qwen2.5 models, is held to a 0.5% cap in every register, not just on average. The current version called 0 of 4,546 verified human documents AI (95% upper bound 0.08%). The register you write in shapes your actual risk.
Every AI detector reports an accuracy figure. The number that better predicts whether your work gets flagged is the false-positive rate for the kind of writing you produce, and the two can tell very different stories. In our own testing, one version of our detector stayed under a 0.5% cap on pre-2021 Medium posts (0.14%) while calling 1.26% of US federal agency documents AI.
The AEO Content AI Content Detector is built on a ModernBERT classifier plus a separate predictability check from a pair of open Qwen2.5 models. It scores passages of about 300 words. A document is called AI when a passage scores 0.998 or higher, human when no passage reaches 0.5, and anything in between is not called. The not-called zone is not a flaw: it is where the detector declines to give a verdict rather than force one.
This article applies one framework throughout: a two-axis check that separates the false-positive rate from overall accuracy. A detector can look strong on one axis while failing on the other, and a single headline figure does not show which.
Hands-on tests show why the second axis matters. In The Best AI Detector Tools - I Tested 11 Tools for False Positives, Ilam Padmanabhan ran a travel blog post he wrote in 2021, before ChatGPT, through a set of detectors. Two well-known detectors flagged his human sample as 65-80% AI-generated, while Winston AI scored it 99% human. That is one writer and one sample, not a benchmark, but it shows how far tools can disagree on the same human text.
The sections below cover where false positives climbed in our tests, how retraining brought agency prose back under the cap, and how we confirmed the result on data that played no part in choosing the model. Any buyer evaluating a detector should ask for the same information: the false-positive rate by type of writing, measured on data that played no part in training.
AI detectors are now used to make consequential calls on published content, student submissions and agency deliverables. Accuracy is the headline number. The false-positive rate for your kind of writing is what decides whether the tool is safe to use on your content.
A false positive is any verified human document a detector calls AI. The rate can vary by type of writing, by passage length and by how closely the documents used to set the threshold resemble the content being tested. A tool that is accurate on a balanced test set may still flag a larger share of human text in kinds of writing it was not tested on. That is not a labeling error in the narrow sense; it is a consequence of where the threshold was placed and what documents were used to set it.
The stakes are rising. According to AI Detector: The Complete 2026 Guide by Mohab A. Karim, the Academic Integrity Council found that 87% of universities updated their AI policies between January 2025 and early 2026, and a RAND survey found 54% of students and 53% of teachers using AI for schoolwork, while only 45% of principals reported having any school or district AI policy. The AEO Content AI Content Detector uses a ModernBERT classifier and a predictability check from a pair of open Qwen2.5 models, held to a 0.5% false-positive cap in every register rather than a single aggregate rate.
What will determine which AI detectors remain credible in the next 12-24 months?
We expect detectors to be judged more on false-positive rates published for each kind of writing than on headline accuracy. These are forecasts from current signals, not findings.
Three signals are already visible in our sources, and all three push the same way: toward buyers who ask for specific error rates rather than one overall accuracy figure.
| Signal | Prediction (12-24 months) | Weak signal today | Why it matters |
|---|---|---|---|
| Detector split by writing type | Buyers shift toward vendors publishing false-positive rates by type of writing, and tools that report only overall accuracy face harder questions in procurement. | Mohab A. Karim's guide reports that in the RAID benchmark, ZeroGPT's false-positive rate plateaued at 16.9% even at its most conservative setting, while Originality.ai held false positives as low as 0.62% under the same conditions. | A single accuracy claim does not show what false-positive rate a tool was tuned to. The error rate a writer experiences depends on the tool and on the kind of writing submitted. |
| Policy specificity increases | Institutions add explicit false-positive ceilings to AI detection policies, moving from "we use a detector" to "the detector must hold below X% for the writing it screens." | The same guide reports that the Academic Integrity Council found 87% of universities updated their AI policies between January 2025 and early 2026, and that in a RAND survey 54% of students and 53% of teachers used AI for schoolwork (both up more than 15 percentage points) while only 45% of principals reported any school or district AI policy. | Anyone whose work is screened depends on whether the deciding institution specifies a ceiling or accepts whatever the vendor claims. A shift toward specificity favors tools that document their rates. |
| Blanket dismissal weakens, individual accusations stay contested | The argument that all detectors are unreliable loses ground as vendors defend sub-1% rates, while using a single score against one writer stays disputed. | Pangram claims a false-positive rate of 1 in 10,000, and Turnitin publishes under 1% false positives on documents with more than 20% AI text. Tim Requarth counters that Pangram "is probably a decent tool for studying population trends," but that using it against individuals is a fundamentally different application. | The debate shifts from whether detection works at all to what a score can support: population trends, or a decision about one person. |
One gap remains. As Karim puts it, the industry is "full of vendors publishing their own accuracy numbers, which is a bit like a company grading its own exam." Edited and paraphrased text is also harder: RAID (ACL 2024) found detectors easily fooled by adversarial attacks, sampling changes, repetition penalties and unseen models, and in our own tests GPT-4 text rewritten sentence by sentence with the DIPPER paraphraser was called AI only 31% of the time by our current detector. Independent, per-writing-type testing is the audit I would want to see before placing consequential trust in any single tool's aggregate claim.
Forward Signal: 12-24 months horizon
Where AI Detection Accuracy Is Headed
Three scored forecasts on how false-positive rates in AI writing detection split by tool and writing population over the next two years.
Forecasts for detector false-positive rates
Read each forecast alongside its confidence to judge which detection claims will hold up as you choose a tool.
The claim that detectors are inherently untrustworthy will weaken as leading tools defend sub-1% rates, with Pangram reporting 1 in 10,000 false positives and Turnitin publishing under 1% on documents more than 20% AI, reframing false positives as a calibration gap between vendors rather than a fatal flaw of the whole category.
Over the next 12-24 months, detector performance will separate sharply by writing type, with leaders like Originality.ai holding false positives as low as 0.62% in the RAID benchmark while tools such as ZeroGPT stay stuck around 16.9%, pushing buyers toward vendors that report rates per register instead of one blended figure.
As 87% of universities have already revised AI policies between January 2025 and early 2026 and student and teacher AI use crosses roughly 54% and 53%, institutions will increasingly require documented sub-1% false-positive thresholds and move away from acting on a single top-line percentage.
Faint signals worth tracking: The RAID benchmark already shows ZeroGPT plateauing at a 16.9% false-positive rate even at its most conservative setting, while Originality.ai holds as low as 0.62% under identical conditions. The Academic Integrity Council found 87% of universities updated AI policies between January 2025 and early 2026, while RAND shows student and teacher use rising more than 15 percentage points. Pangram claims a false-positive rate of 1 in 10,000, and a University of Maryland preprint co-authored by four Pangram employees used its detector to scan 250,000 newspaper articles, though that scan estimated AI text in the news rather than testing false positives on human writing.
Supporting and contrary detector data
Each forecast lists both the benchmarks that back it and the studies that push the other way.
- AI Detector: The Complete 2026 Guide - by Mohab A.Karim - Substack points the same way. [Substack / Newsletter]The RAND survey found 54% of students and 53% of teachers reported using AI for schoolwork (both up more than 15 percentage points year-over-year), while only 45% of principals reported having any school or district AI policy. “NPR reported on a 17-year-old student, Ailsa Ostovitz, wrongly accused of academic misconduct after a detector returned a 30.76 percent AI probability score on…”
- The case rests on The Problem with AI Detector Companies - by Tim Requarth. [Substack / Newsletter]Pangram's detector flagged the "Shy Girl" novel that Hachette canceled, and the viral Modern Love column shared on X.
- AI Detector: The Complete 2026 Guide - by Mohab A.Karim - Substack supports this forecast. [Substack / Newsletter]The Academic Integrity Council found 87% of universities updated their AI policies between January 2025 and early 2026.
- AI Detector: The Complete 2026 Guide - by Mohab A.Karim - Substack is what puts this forecast on the board. [Substack / Newsletter]In the RAID benchmark, ZeroGPT's false-positive rate plateaued at 16.9% even at its most conservative setting, while Originality.ai held false positives as low as 0.62% under the same conditions.
What could shift these forecasts
Real-world conditions, from non-native-writer error rates to base-rate math, that would reverse the outlook below.
Before you rely on these numbers
A score measures how much current evidence backs a call, and that evidence keeps moving. The top forecast here sits at 50/100, while the minority view at 50/100 shows where the sources still disagree.
- Top detectors become harder to dismiss. That call weakens first if regulators or buyers move in the opposite direction.
- Top detectors become harder to dismiss. That one becomes the more durable forecast if the source mix shifts toward stronger contrary evidence.
Where does the false-positive rate climb, and which writing register is most exposed?
Formal institutional prose. Our previous detector called 8 of 637 US federal agency documents AI (1.26%) and 4 of 2,868 pre-2021 Medium posts (0.14%), seven times the 0.02% we saw on marketing blogs.
The Medium result is worth reading closely. The rate there was 0.14%, with a 95% upper bound of 0.36%, under our cap overall. The four posts called AI were a USAID program update, a policy analysis of an Affordable Care Act bill, a technology company's design case study and a list of 100 headline-style post titles. Three of the four were neutral, well-organized institutional prose, and there were no false positives in tech tutorials, data science, crypto and finance, startups, personal essays or culture writing. The agency documents pointed the same way: the hardest source was the State Department's ShareAmerica, plain English written for international readers, with 5 of its 90 documents called AI.
Our data does not show why institutional prose gets misread. A plausible explanation is that it is the register AI models imitate by default: neutral, explanatory and information-dense. Reinhart et al. (PNAS, February 2025) found model writing noun-heavy and information-dense, with a larger gap from human writing for instruction-tuned models than for base models. We did not test why marketing blogs came out lower, so we do not offer a reason.
A common assumption is that older human content is safe from AI detectors. The Medium archive predates ChatGPT by years, yet the false-positive rate on it was still seven times the rate on marketing blogs and client archives. Age did not protect the text. Dates also mislead in the other direction: in one payments company's blog, 7 of 50 posts from January to October 2022 already read like GPT-3 or Jasper output, so we only trust human labels for text dated 2021 or earlier.
Our archives are not the only place the rate climbs. As James Kuhman summarizes in What an AI Detector Cannot Prove, Liang and coauthors tested seven detectors in 2023 on 91 TOEFL essays by non-native English speakers: on average the detectors falsely classified 61.22% of them as AI-generated, all seven misclassified the same 18 essays, and the same detectors performed nearly perfectly on 88 essays by American eighth-grade students.
Passage length is a separate variable. When we cut verified human Medium posts into 50-, 100-, 200- and 300-word chunks, our detector called 1.0% of the 50-word chunks AI and none of the longer ones. Very short tails caused trouble elsewhere: two of three false alarms in one 2019 blog were 47- and 62-word call-to-action tails, so we now drop tails under 150 words. A plausible explanation is that very short passages give a classifier too little text to go on, but we have not tested that directly.
Training runs add a third source of variation. Retraining the identical recipe moved the not-called rate on human documents from 0.22% to 0.66%, and recall on pipeline articles at the strict threshold from 97.8% to 92.4%. The runs varied as much as the effects we were trying to measure, so no single run is canonical. We only ship a model after confirming it on fresh data that took no part in choosing it.
Small rates also add up with repetition. In The Problem with AI Detector Companies, Tim Requarth cites Princeton computer scientist Arvind Narayanan's calculation that if every instructor ran an AI detector on all student work, roughly 500 to 1,000 written works over four years, 5 to 10 percent of the student body would be falsely accused of cheating at some point. A rate that looks tiny per document stops looking tiny across a whole degree.
The not-called zone, documents whose highest passage scores at least 0.5 but below 0.998, is our response to that uncertainty. Rather than force a verdict, the detector withholds one. On pre-2021 Medium posts, 0.73% were not called, against 0.22% on the older pool: coverage we give up rather than force a call.
How does targeted recalibration reduce agency prose false positives, and what does confirmation on fresh data prove?
By retraining: we added 100 documents from other federal agencies to the training set. Agency false positives fell from 8 of 637 (1.26%) to 2, with none on 511 documents from unseen agencies.
The previous detector (v8d) was above our 0.5% cap on agency prose. The new training documents came from agencies such as the CDC, NIH, FDA and EPA. We keep whole sources on one side: an agency is used either to train the detector or to test it, never both. Our agency test used 9 agencies for testing, 11 for training, and 8 more that we collected only after choosing the model. The retrained version (v9b) called 2 of the 637 agency documents AI, 0 of 4,546 verified human documents and 2 of 2,868 Medium posts, while still catching 99.6% of held-out AI articles.
Because we trained and compared five candidate models, picking the best one on those test sets could flatter it. So we confirmed the winner on data that took no part in choosing it: 511 documents from 8 agencies never used anywhere, where the old detector made 3 false positives and the new one 0, and 1,401 fresh Medium posts from authors not in the first sample, again 3 before and 0 after.
Earlier versions show how much one training cycle can change. Before we recalibrated it, our v6 head called 2.06% of a 2,931-document human pool AI, far above the cap. Version v7 called 1 of 4,546 (0.02%), and the current version 0 of 4,546 (95% upper bound 0.08%). Improvements did not spread evenly: v8d was under the cap on Medium posts while breaking it on agency prose, which is exactly the kind of concentration an average hides.
On pre-2021 Medium posts, the not-called rate rose to 0.73%, against 0.22% on the older pool. Documents in the not-called zone receive no verdict: the detector withholds a call rather than force one. We do not know why Medium posts landed there more often; the data records that they did, not the reason.
The published numbers carry limits. They are internal evaluations on held-out data, not an independent audit. The language models used for scoring have likely read old public text during their own training, which can make human text look more predictable to them. And the strongest proof, human writing that was never published anywhere, is the next dataset we are adding. Until then, the confirmation sets are the best evidence we have that the rate holds on text the model never saw.
Outside tools face the same test of consistency. In The Problem with AI Detector Companies, Tim Requarth reports that the Wall Street Journal's James Taranto ran three op-eds through Pangram, which claims a false-positive rate of 1 in 10,000, and got results that did not match the researchers' own labels; Taranto called them "wildly inconsistent." Requarth later conceded that Pangram is "quite accurate on fully AI-generated text from major models, even on short bits of text." A vendor's claimed rate is a starting point for questions, not a substitute for results you can check.
The clearest conclusion from both versions is that a single average hides where errors concentrate. The 1.26% agency rate was not spread across all writing; it sat in one register, and adding training examples of that register brought it back under the cap. A cap checked in every register exists to catch exactly that kind of hidden concentration before it reaches a specific writer's work.
What does false-positive rate actually measure, and why doesn't detector accuracy tell you?
A false-positive rate is the share of verified human documents a detector calls AI. Our current detector called 0 of 4,546 such documents AI, with a 95% upper bound of 0.08%.
Accuracy and false-positive rate are separate numbers, tied together by the threshold. Mohab A. Karim's guide sums up a point the RAID benchmark makes: you cannot fairly compare two detectors' accuracy numbers unless you know what false-positive rate each one was tuned to. A detector tuned to catch more AI text will usually flag more human writing, and one held below a fixed false-positive ceiling has to accept some misses. We saw that trade-off in our own training: adding 404 human marketing posts drove human false positives to zero but cost 13-27 points of recall on AI articles.
Accuracy is the share of all documents classified correctly, and it can look strong while false accusations pile up. James Kuhman illustrates the base-rate problem in What an AI Detector Cannot Prove (Medium, September 2026) with a hypothetical: among 1,000 submissions in which 50 break an AI-use rule, a detector that finds 40 of them but falsely flags 5% of the other 950 produces about 48 false accusations alongside 40 correct flags. By our arithmetic that detector is still about 94% accurate, yet most of the work it flags is human.
A useful frame for evaluating any detector is what I call the two-axis check. First: what share of AI text does the detector catch? Second: what share of verified human text in your specific writing register does it flag? The second question determines whether the tool is safe to use on your content.
A common misconception is that high overall accuracy implies a low false-positive rate on human writing. It does not, and careful writers can pay for the gap. A short explainer video, Turnitin AI Detection False Positives | Can AI Flag You By Mistake, warns: "if your writing is too perfect too structured or too clean it can get flagged even if it is 100% human," because skilled human writing and AI writing are starting to look similar. That is the narrator's explanation rather than a measured result, but it fits what we saw: the human writing our previous detector misread most often was neutral institutional prose.
Our detector is held to a 0.5% false-positive cap in every writing register, not just on average. The rule applies to each kind of writing we test, from client archives dated 2021 or earlier and archived marketing and SEO blogs to open-web text, sports and health copy, pre-2021 Medium posts and US federal agency prose. The current version called none of 4,546 documents across 17 registers AI, which carries a 95% upper bound of 0.08%: with 95% confidence, the true rate on writing like that pool is no higher than 0.08%. Across all 9,963 verified human documents we tested, it called 4 AI.
Specificity of evidence may also shape how a passage reads, though this result comes from AI-drafted articles, not human writing. On one sports equipment manufacturer's site, sections built on the company's own sales and shipping records read AI 37% of the time, against 71% for the other sections of the same articles. For a payments company whose articles used general industry figures, the split was 98% against 100%, so the effect is not general. A plausible explanation is that facts only one company could publish are harder to write generically, but the data does not prove it.
The question most buyers ask is whether an AI detector is accurate. The more useful question is whether its false-positive rate holds below a documented ceiling in every kind of writing it screens. Those are different questions, and the answer to the first tells you little about the second.
Some vendors now publish low headline false-positive rates: Pangram claims 1 in 10,000, and Turnitin publishes under 1% on documents with more than 20% AI text. Treat those as claims to check, not results. The question that belongs in any evaluation is not "how accurate is this tool?" but "what is the false-positive rate on the kind of writing I produce, and on what documents was it confirmed?"
The right response is not to dismiss detectors as unreliable. It is to require a documented error rate by type of writing from whoever runs them, and to check that the confirmation data was held out rather than drawn from the pool used to choose the model. Our own numbers are stated above: 0 of 4,546 verified human documents called AI and 4 across all 9,963, internal results rather than an independent audit. You can test a draft with the free AI Content Detector.
We apply the same standard of evidence to writing. The AEO Content Engine produces articles grounded in each company's own records, written to read like expert human writing and structured so AI answer engines can cite them.
This article is part of our research series on how AI writes and how humans write. The overview of the whole series is How AI Writes vs How Humans Write.
Written by
Alex Shortov
CTO, AEO Content
Full-stack engineer and content infrastructure architect with 20 years of building enterprise systems.
Connect on LinkedInSummarize This Article With AI
Open this article in your preferred AI engine for an instant summary.
Frequently Asked Questions
What is a false positive in AI detection?
A false positive occurs when a detector labels a verified human document as AI-generated. It is the error that matters most for anyone submitting genuine work. A trustworthy rate is measured on documents that played no part in choosing the model; ours was confirmed on 1,401 fresh Medium posts and 511 documents from agencies never used anywhere, with 0 false positives on each.
What is the "not-called" zone, and why does it exist?
Our detector withholds a verdict when a document's highest-scoring passage lands at 0.5 or above but below 0.998. If no passage reaches 0.5 the document is called human, and if any passage scores 0.998 or higher it is called AI. The zone keeps uncertain documents from being forced into a label. The cost is coverage: 0.22% of documents in our older human pool were not called, and 0.73% of pre-2021 Medium posts.
Does the writing register I use change my false-positive risk?
Yes, it can. Our previous detector called 0.14% of pre-2021 Medium posts AI but 1.26% of US federal agency documents, roughly nine times as high. That is why a cap has to hold independently in each kind of writing, not only on average.
How should detection policies account for false positives?
As guidance rather than a finding: a policy that acts on detector output should name an acceptable false-positive ceiling, say what kind of writing it was measured on, and require corroboration before a penalty. Policies are changing quickly; according to Mohab A. Karim's guide, the Academic Integrity Council found 87% of universities updated their AI policies between January 2025 and early 2026. Base rates matter too: in James Kuhman's hypothetical, a detector that falsely flags 5% of 950 honest submissions produces about 48 false accusations alongside 40 correct flags.
Why doesn't overall accuracy tell me what I need to know?
Accuracy combines true positives and true negatives into one number. That number can look strong even when a detector is producing a high false-positive rate on a specific writing population. The false-positive rate in the register you actually write in is the figure that predicts your individual risk.