A detector returns 94 percent AI-generated. A student is called into an office. Somewhere in that chain, a probability estimate turned into an accusation, and almost nobody involved could explain how.

The belief underneath the practice is that AI text detection is basically solved. Paste in the essay, read the number, act on it. The evidence does not support that, and the way it fails is worse than simple inaccuracy.

The tools flag good writing

The most uncomfortable finding in this area is that detectors do not fail randomly. They fail in a direction.

Research summarised by the Authors Guild describes what it called a troubling paradox: the more refined and controlled a writer's style, the more closely it resembles the output these tools are trained to catch. Clean structure, consistent register, careful transitions. Those are the marks of a practised writer and also the statistical signature of a language model.

The people most likely to be falsely accused, then, are careful writers. That includes second-language writers taught to follow explicit structural rules, and students who have been drilled in exactly the essay conventions that now read as machine-like.

The problem may not be an engineering problem

A paper published in March 2026 made a stronger claim than "current detectors need work". Working from formal probability rather than benchmark testing, the authors argued that any text-only, one-shot detector with meaningful detection power will necessarily produce false accusations among writers whose style overlaps statistically with model output.

If that argument holds, the false positive rate is not a bug awaiting a better classifier. It is a ceiling created by the diversity of human writing itself. I would treat this as the most important open question in the field, and I want to be clear that it is a theoretical result rather than a settled one.

Detection also fails in the other direction

The tools are not uniformly strict either. Published comparisons show GPTZero losing much of its detection ability once text is run through a humanising tool, with false negative rates around 50 percent and above across most genres and models tested. Pangram holds up considerably better against the same modifications.

So the picture is not "detectors are too aggressive". It is that they can be too aggressive toward honest careful writers and too permissive toward anyone who spends thirty seconds running their output through a paraphraser. Both errors land on the wrong people.

This is a familiar shape. Meta's own image detector missed half of its own AI-generated images once they were cropped. Detection systems tend to be brittle in exactly the situations where somebody is actively trying to defeat them.

What the number actually means

A detector score is a statement about how closely a passage resembles patterns in the tool's training data. It is not a record of what happened when the text was written. Those are different claims, and only one of them is evidence.

Used as a prompt for a conversation, a detector score is fine. Used as proof, it asks a probability estimate to do a job it was never built for. The honest version of the tool's output would read something like "this writing is stylistically regular", which is a much less useful thing to put in a disciplinary file.

Sources

  1. i. www.techtimes.com
  2. ii. bfi.uchicago.edu
  3. iii. gptzero.me
  4. iv. arxiv.org

Commentarii · 0

Add · a · Comment