Ask most people whether a teacher or an editor can tell that AI wrote something, and they will point to a detector. Paste the text in, read the percentage, case closed. The tools are marketed with exactly that confidence, and schools and workplaces have bought in. The evidence for how well they actually work is a good deal weaker than the sales pitch.

The clearest admission came from OpenAI itself. In January 2023 the company launched a classifier to spot AI-written text. Six months later, on July 20, 2023, it quietly shut the tool down, citing a low rate of accuracy. Its own numbers were poor: the classifier correctly flagged just 26 percent of AI-written text while wrongly labeling 9 percent of human writing as machine-made.

False positives are the real problem

A detector that misses AI text is an annoyance. One that accuses innocent people is a different kind of failure, and it is the more common one. Vanderbilt University disabled Turnitin's AI detection in 2023 after doing the arithmetic: even a 1 percent false-positive rate, applied across the 75,000 papers its students submit, would mean roughly 750 wrongful accusations a year.

The errors are not spread evenly. One widely cited Stanford study found that seven detectors flagged essays by non-native English speakers as AI-generated about 61 percent of the time, while almost never misfiring on writing by native speakers. Students who lean on simpler vocabulary and repeated phrasing, including many neurodivergent writers, get caught in the same net. The tools mistake predictable prose for machine prose.

What the tools can and cannot do

Detectors are least reliable exactly where they get used most, on short assignments. Under a few hundred words there is too little signal for an honest estimate, yet that is the length of a typical homework paragraph. Turnitin now advises against putting weight on scores in its lower range for this reason.

None of this means AI writing is undetectable, or that concern about it is imaginary. It means a detector score is a weak signal, not a verdict. Treating a percentage as proof, with a grade or a job on the line, asks far more of these tools than they can deliver. The honest position is the uncomfortable one: for now there is no reliable way to prove that a specific piece of writing came from a machine.

Sources

  1. i. techcrunch.com
  2. ii. lawlibguides.sandiego.edu
  3. iii. www.turnitin.com
  4. iv. pmc.ncbi.nlm.nih.gov
  5. v. www.popularai.org

Commentarii · 0

Add · a · Comment