What we tested
The benchmark has three kinds of text, all between 40 and 800 words:
- 1,260 human texts from 44 sources: news (BBC, CNN, CC-News), student and learner essays (PERSUADE, TOEFL, ELLIPSE, Chinese English learners), parliamentary and UN speeches, PubMed abstracts, technical docs, reviews, forum posts and web pages in seven languages. Learner essays are included on purpose: they are where detectors most often accuse real students.
- 891 AI texts from 62 models by 11 vendors, including GPT-6 (Sol, Luna, Astra), Claude Opus 5.5 and Fable 5.1, Gemini 3.8, Grok 4.7, DeepSeek V4.1, Qwen 3.8 Max, GLM 5.3 and MiMo. 150 of them were paraphrased by a second model.
- 421 humanized texts: AI text run through StealthGPT, CleverHumanizer, WriteHuman, WalterWrites, UndetectedGPT and Rephrasy.
ZeroGPT and our models scored every text. Pangram scored 513 and GPTZero 894, limited by credits. Each comparison below only counts texts that every listed detector scored.
How we count a detection. Third-party tools count as a hit when they return "AI". For tropa-3 we use a display score of 90 or more, the threshold that flags 0.1% of 239,490 human texts none of our models ever saw in training. tropa-2 is held to the same standard: its threshold is set so it also flags 0.1% of those texts. We report "Mixed" verdicts separately.
Results on the shared set
| 513 texts | n | ZeroGPT | Pangram | GPTZero | tropa-2 | tropa-3 |
|---|---|---|---|---|---|---|
| Human flagged as AI ↓ | 247 | 0.8% | 0.0% | 0.0% | 0.4% | 0.0% |
| Raw AI detected | 159 | 21.4% | 44.0% | 61.6% | 77.4% | 95.6% |
| Paraphrased AI | 33 | 9.1% | 9.1% | 6.1% | 18.2% | 54.5% |
| Humanized AI | 74 | 4.1% | 35.1% | 25.7% | 28.4% | 54.1% |
The gap on raw AI text is the headline. Pangram and GPTZero return "AI" for fewer than two thirds of texts from current models while flagging no humans. tropa-3 catches 96% and also flags no humans. The 95% confidence interval for tropa-3 is 91–98%, for GPTZero 54–69% and for Pangram 37–52%, so the difference is not noise.
"Mixed" is where the other detectors hedge
All three third-party tools have a middle verdict. If you count "Mixed" as a hit, Pangram reaches 81% on raw AI and GPTZero 84%, still with zero human texts flagged. ZeroGPT reaches 52% but then marks 19% of human texts as partly AI. Our site has a middle band too: "Likely AI" for scores from 70 to 89. Counted the same way, tropa-3 stays ahead in every row.
| "AI" or "Mixed" counted | n | ZeroGPT | Pangram | GPTZero | tropa-3 (70+) |
|---|---|---|---|---|---|
| Human flagged ↓ | 247 | 19.0% | 0.0% | 0.0% | 1.2% |
| Raw AI detected | 159 | 52.2% | 81.1% | 84.3% | 98.1% |
| Paraphrased AI | 33 | 39.4% | 54.5% | 48.5% | 63.6% |
| Humanized AI | 74 | 29.7% | 50.0% | 27.0% | 70.3% |
tropa-3 column: score 70 or more, the "Likely AI" and "AI-generated" bands on our site. The price is 3 of 247 human texts flagged as likely AI.
A larger sample without Pangram
With 894 texts the paraphrase and humanizer groups are large enough to trust. The picture holds.
| 894 texts | n | ZeroGPT | GPTZero | tropa-2 | tropa-3 |
|---|---|---|---|---|---|
| Human flagged as AI ↓ | 252 | 0.8% | 0.0% | 0.4% | 0.0% |
| Raw AI detected | 207 | 23.2% | 60.4% | 75.8% | 95.2% (91–97) |
| Paraphrased AI | 149 | 5.4% | 12.1% | 30.9% | 61.7% (54–69) |
| Humanized AI | 286 | 1.7% | 24.1% | 31.5% | 57.7% (52–63) |
Brackets show the 95% confidence interval for tropa-3. tropa-2 is shown at the same 0.1% false-positive rate as tropa-3.
Humanizers, one by one
Humanizers are built to beat detectors, and it shows. StealthGPT is caught by GPTZero and by us. The others slip past GPTZero almost completely. UndetectedGPT and WalterWrites are the hardest cases for our model too.
| Humanizer | n | ZeroGPT | GPTZero | tropa-2 | tropa-3 |
|---|---|---|---|---|---|
| StealthGPT | 49 | 0.0% | 87.8% | 10.2% | 87.8% |
| Rephrasy | 48 | 10.4% | 27.1% | 50.0% | 77.1% |
| WriteHuman | 45 | 0.0% | 8.9% | 57.8% | 75.6% |
| CleverHumanizer | 48 | 0.0% | 6.2% | 41.7% | 50.0% |
| WalterWrites | 48 | 0.0% | 12.5% | 22.9% | 31.2% |
| UndetectedGPT | 47 | 0.0% | 0.0% | 8.5% | 25.5% |
At the same false-positive rate, tropa-3 beats tropa-2 on every humanizer, most clearly on StealthGPT (87.8% vs 10.2%). If you accept more false positives, both catch more: at a looser threshold of 70, tropa-3 catches 71% of humanized text and flags 1.6% of humans in this set.
False positives at scale
247 human texts can show that a detector is careful. They can't show how rare its mistakes are. For that we score 239,490 human texts that none of our models saw in training: web pages in seven languages, news, speeches, biomedical papers, docs, e-mails, forums and reviews.
| 239,490 human texts | tropa-2 | tropa-3 |
|---|---|---|
| Flagged at the live threshold | 3.02% | 0.10% |
| Flagged when both catch 97% of raw AI | 1.99% | 0.32% |
| … news only (10,591) | 6.07% | 0.30% |
| … speeches only (14,000) | 4.01% | 0.22% |
| … biomedical only (16,280) | 2.63% | 0.35% |
That is a 30-fold drop in wrongly flagged human writing between our last two versions. We did not run the third-party tools on this set.
Non-native English writers
In 2023, Liang et al. showed that AI detectors flagged most TOEFL essays written by non-native English speakers as AI. Those essays are part of this benchmark, and they were never part of our training data. tropa-3 flagged none of the 60 (median score 7 of 100). ZeroGPT labelled none of them AI but marked 15 of 60 as mixed. Pangram and GPTZero, which scored 14 of them, flagged none.
Limits of this benchmark
We tested on our home turf. The AI and humanized texts come from the same models and humanizers our detector was trained on. The prompts and documents are different and no text overlaps, but the other detectors did not get to train on this mix. Read our humanizer numbers with that in mind. The TOEFL essays above are the exception: they come from a published study and were never in our training data.
- Pangram scored 513 texts, GPTZero 894. Per-humanizer numbers for Pangram are too small to report.
- Third-party results were collected on 24–25 September 2026 through the public web apps (Pangram, GPTZero) and ZeroGPT's web API. These tools update often.
- tropa-3 reads the first 512 tokens of each text. All benchmark texts are 800 words or fewer.
- "Human" means written before or independently of LLMs as far as the source guarantees. A few web texts may contain AI text we could not identify.
Get the data
Every text, label and verdict is on Hugging Face, including every row where our model is wrong. 1,595 texts are included in full: all AI and humanized texts plus human texts under open licences. For copyrighted human texts we publish the source dataset, original id and URL, plus a SHA-256 hash so you can confirm you fetched the exact text. The same snapshot is in the GitHub repo under 2026-09/.
from datasets import load_dataset
ds = load_dataset("wasitaigeneratedcom/ai-detector-benchmark-2026-09", split="train")
df = ds.to_pandas()
shared = df[df.in_pangram_subset]
print(shared.groupby("group")[["pangram_label", "gptzero_label"]].value_counts())Columns include zerogpt_label, pangram_label, gptzero_label, tropa3_score and tropa3_display. The dataset card documents every field and licence.
FAQ
Pangram vs GPTZero vs ZeroGPT vs tropa-3: which AI detector is more accurate in 2026?
On the 513 texts all five detectors scored, tropa-3 detected 95.6% of raw AI text, GPTZero 61.6%, Pangram 44.0% and ZeroGPT 21.4%. tropa-3, Pangram and GPTZero flagged none of the 247 human texts; ZeroGPT flagged 0.8%. The AI and humanized texts come from the same models and humanizers our detector was trained on, so treat our numbers as a home-field result.
Is Pangram accurate?
Pangram flagged none of the 247 human texts. It returned "AI" for 44.0% of raw text from current models. If its "Mixed" verdict counts as a hit, it reaches 81.1% on raw AI text, still with no human text flagged.
Does counting "Mixed" as AI change the ranking?
It narrows the gap but does not change the order. With "Mixed" counted, Pangram reaches 81.1% and GPTZero 84.3% on raw AI text, both with zero human texts flagged. Counting our own middle band ("Likely AI", score 70 or more) the same way, tropa-3 reaches 98.1% on raw AI and flags 1.2% of human texts. ZeroGPT reaches 52.2% but marks 19.0% of human texts as partly AI.
Can AI detectors catch humanized text?
Only partly. On 286 humanized texts, tropa-3 caught 57.7%, GPTZero 24.1% and ZeroGPT 1.7%. StealthGPT is caught by GPTZero and by tropa-3 (87.8% each). UndetectedGPT and WalterWrites are the hardest cases: tropa-3 caught 25.5% and 31.2%, GPTZero 0.0% and 12.5%.
How does tropa-3 compare with tropa-2?
We compare both at the same false-positive rate: the threshold at which each model flags 0.1% of 239,490 held-out human texts. At that rate tropa-3 detects 95.2% of raw AI text in the 894-text set against 75.8% for tropa-2, 61.7% of paraphrased text against 30.9%, and 57.7% of humanized text against 31.5%. At its old production threshold tropa-2 catches more, but it then flags 3.02% of the held-out human texts.
How often does tropa-3 flag human writing as AI?
It flagged 0 of 247 human texts in the shared set and 0.10% of 239,490 held-out human texts that none of our models saw in training. On news it flagged 0.30%, on speeches 0.22% and on biomedical papers 0.35%. No detector verdict should be the sole evidence in an academic integrity case.
Check a document yourself
Paste it into WasItAIGenerated. You get the same tropa-3 score used in this post.
Test tropa-3 Free