Say it plainly
An "AI score" may mean nothing for the genre you write in
August 2026. Claude Opus 5, GPT-5.6, DeepSeek V4, GLM-5.3 and Kimi K3 each produced
50–100 samples per genre, set against 50 public human texts, recomputed
after matching length on both sides.
| Genre | Separation | Verdict |
| Short explainer / news | 0.837 | still useful |
| Social posts | 0.771 | usable |
| Academic abstracts | 0.621 | already weak |
| Fiction | 0.479 | a coin flip |
| Long-form blogs | 0.155 | inverted |
Read it as: pick one machine text and one human text at random — this is the probability the
machine one scores higher. The tick is 0.5, a coin flip. Below that, the
direction is reversed: asked to pick "the one that looks AI-written", it picks the human more
often. The long-form row has a thinner sample (15 machine vs 22 human) than the others, but
the direction is unambiguous.
Human-written papers get flagged too
60 real CNKI abstracts published before large language models existed average 60/100, and 35% land in the "almost certainly AI" band. Chinese academic prose is full of set connectives; the detector cannot tell convention from generation.
So don't screen long text with it
Short expository writing is a genuine exception — don't throw that out. But using it to judge whether a long blog post or a novel chapter was machine-written gives you answers that are wrong in a systematic direction.