Three days after OpenAI released GPT-6 Astra, Originality.ai published a 1,000-response benchmark reporting 99.9% detection. We rebuilt their design and ran it through tropa-2. The headline is quickly told — and it is the least interesting part of this page.
On 5 September, one day after generating the samples, Originality.ai reported that its detector flagged 999 of 1,000 English GPT-6 Astra responses as likely AI — 200 in each of five categories: blog, conversational, expository, news and review. The single miss was a conversational response.
The study also states its own limitation, in its own words: “English, known-AI text only; no human control group or false-positive measurement.” We want to be precise about that, because it would be easy and unfair to present it as something they concealed. They disclosed it plainly. It is a scope decision, not an omission — and we made the same one here.
We mirrored the design rather than inventing our own, so that the two results are directly comparable: 200 fresh GPT-6 Astra generations in each of the same five categories, from a fixed matrix of 40 topics crossed with 5 framings, so no prompt repeats within a category. Each response was scored once by tropa-2.
Originality counted a response as detected at 0.50 on their scale. tropa-2 reports 0–100, so 50 is the mirror of that setting. But 50 is more sensitive than what we actually ship: wasitaigenerated.com displays “AI-generated” at 90.
Mixing those two would be the easiest way to flatter ourselves, so we report both. The result was identical at each, which is why the headline above uses the stricter one.
Two further notes, so the numbers can be read for what they are. Astra prepends a Markdown heading to most responses; we left it in place, because that is what real pasted text looks like, and Originality does not state whether they stripped it — it is a plausible source of small divergence between the two studies. And 4 of the intended 1,000 generations failed at the API and were excluded rather than retried into the set, leaving n = 996.
| Category | Detected (≥ 90) | Counts | 95% CI |
|---|---|---|---|
| Blog | 100.0% | 200/200 | 98.1–100.0% |
| Conversational | 100.0% | 200/200 | 98.1–100.0% |
| Expository | 100.0% | 200/200 | 98.1–100.0% |
| News | 100.0% | 196/196 | 98.1–100.0% |
| Review | 100.0% | 200/200 | 98.1–100.0% |
| All categories | 100.0% | 996/996 | 99.6–100.0% |
Every category came back clean, with no weak spot across the five writing styles. We publish the counts and intervals rather than a bare percentage so the result can be checked rather than taken on trust.
This study measures one thing: whether GPT-6 Astra output is detectable. By design it runs only on text known to be machine-written, so it produces a detection rate and nothing else.
The companion question — how a detector behaves on writing by actual people — is one we report separately, because it does not change with each model release. In our published detector benchmark, tropa-2 wrongly flagged 1 of 251 human texts, alongside per-tool comparisons against GPTZero and ZeroGPT on the same set.
Those figures, their sample sizes and the methodology are in our detector benchmark. Read this page alongside that one.
The interesting finding is not that we scored 100%. It is that two detectors built by competitors, using different models and different training data, independently reached the same conclusion about a frontier model three days after its release. A new model no longer buys a window of undetectability.
It also says something about how quickly detection now adapts. A newly released model used to buy a period of uncertainty while detectors caught up. On this evidence, that gap has closed to days.
One thing to hold on to when reading any detection figure, ours included: a score is evidence, not proof. A high number on a single document is a reason to ask a question, never a verdict about a person. If you are making a decision that affects someone, the number belongs in the conversation, not in place of it.
The harness that generated and scored these samples is one script in our repository, the prompt matrix is fixed and documented in it, and the aggregate counts behind every number on this page are published as JSON — this page reads them directly, so the prose cannot drift from the data. tropa-2, the model behind these scores, is hosted, so reproducing the exact figures above means calling the API. If you would rather inspect how this kind of scoring works without going through us, our smaller sibling model tropa-mini is published as open weights under Apache-2.0 on Hugging Face. It is a different, smaller model and will not reproduce the numbers on this page.
In this test, yes — comprehensively. tropa-2 flagged 996 of 996 GPT-6 Astra responses as AI-generated at our live threshold, across five writing categories. Originality.ai reported 99.9% on the same study design. Two independent detectors reaching effectively the same result three days after launch suggests Astra's output carries the same statistical signature as earlier frontier models.
It covers raw, unedited GPT-6 Astra output across five writing categories, scored at our live threshold. A detection rate is measured only on text known to be AI-written, so it answers whether this model is detectable — not how a detector behaves on human writing. We report that separately, in our standing detector benchmark.
Because we publish them separately. This study is deliberately narrow: it answers whether a model released three days ago is detectable. False-positive behaviour does not change per model release, so we report it in our standing detector benchmark, where tropa-2 wrongly flagged 1 of 251 human texts alongside comparisons to GPTZero and ZeroGPT.
We mirrored their design deliberately: 200 responses in each of the same five categories, generated fresh, scored at their reported threshold. They found 999 of 1,000. We found 996 of 996 at that same threshold. This is a replication, not a rebuttal — their study and ours agree.
Both. Originality.ai scored a response as detected at 0.50 on their scale; tropa-2 reports 0–100, so 50 is the mirror. We also report at 90, the threshold wasitaigenerated.com actually uses to display "AI-generated". The result was the same at both, which is why we lead with the stricter one.
The generation side, yes: the harness is a single script, the prompt matrix is fixed and documented in it, and the aggregate counts are published as JSON. The scoring side needs tropa-2, which is a hosted model, so reproducing these exact figures means calling the API. Our open-weights release on Hugging Face is tropa-mini, a smaller model — useful for inspecting how the scoring works, but it will not reproduce the numbers on this page.
Sentence-level results in seconds, on text, images, audio and video — with our error rates published, not hidden.