We took 40 passages we knew were AI-generated, cut them to different lengths, and scored each one. The text never changed — only how much of it the detector saw. It caught everything at 100 words. At 25 it caught 23%.
| Words shown | Median score | Detected | 95% CI |
|---|---|---|---|
| 25 | 59.8 | 23% (9/40) | 12–38% |
| 50 | 95.4 | 73% (29/40) | 57–84% |
| 100 | 99.1 | 100% (40/40) | 91–100% |
| 200 | 99.5 | 100% (40/40) | 91–100% |
“Detected” means a score at or above 90, the threshold this site uses to display “AI-generated”. At the looser 50 threshold the same passages reach 65% at 25 words and 93% at 50.
Our documentation said accuracy suffers below 50 words. The measurement says the floor is 100. At 50 words, better than a quarter of text we knew to be machine-written came back below the threshold we display — and a reader would have taken that as a clean result.
That matters because short passages are the ordinary case. People check a suspicious paragraph, a product review, the opening of a cover letter. Those are rarely 100 words, and a false negative there is the worse failure: nothing about a low score invites a second look.
We have corrected the figure in the docs and in the MCP tool description, and passages below the floor now return a low-confidence flag rather than a bare number. We would rather publish this than have someone discover it against their own text.
Below the floor the detector under-reports. A short passage that scores low is weak evidence, not a verdict of human authorship; a short passage that scores high still means something. The asymmetry is worth stating plainly, because it runs toward missing AI text rather than toward accusing a human writer — which is the error we care most about avoiding. Our standing false-positive figures are in the detector benchmark.
40 GPT-6 Astra generations, eight from each of five writing categories, drawn from the set used in our Astra replication. Each was truncated from the start to 25, 50, 100 and 200 words and scored separately, so every row above is the same 40 texts. Counts are published as JSON and this page computes the percentages and Wilson intervals from them at render time.
Limits: one generator, one truncation strategy, and passages cut down rather than written short — a genuinely brief text may behave differently. We have not yet measured false positives at these lengths, and we should: if length moves the score this much, it moves it in both directions.
On our own detector, about 100 words. At that length it flagged all 40 known-AI passages we tested. At 50 words it caught 73% of the same text at our display threshold, and at 25 words only 23%. Shorter passages give the detector less signal, and the score drops accordingly — on text that has not changed at all.
Detection reads statistical patterns across a passage rather than judging individual words. The fewer words there are, the less pattern there is to read, and the estimate falls back toward the middle. The text is no less machine-written; there is simply less of it to measure.
We had published 50 words as the point below which accuracy suffers. The measurement says 100. We have corrected the figure in the API documentation and the MCP tool description, and short passages now come back flagged as low-confidence rather than answered silently.
No — it means a low score on a short passage is weak evidence rather than a clean bill of health. A high score on a short passage still means something. The asymmetry matters: below the floor we under-report, so the error runs toward missing AI text rather than falsely accusing human writing.
100 words or more, and you get a result worth acting on.
Try the Detector Free