OpenAI announced yesterday that GPT-5.5 outperforms humans on several benchmarks. Earlier this year, its GPT-5.4 model crossed the human baseline on OSWorld, a test of desktop application control, scoring 75.0% against a human reference of 72.4%. The headlines write themselves: AI surpasses human performance, models smarter than people. It is worth pausing on what those claims actually mean, because the gap between benchmark scores and general intelligence is wider than the coverage usually suggests.

What a benchmark actually measures

A benchmark is a narrow, predefined task. OSWorld tests whether a model can operate software, fill forms, and navigate browsers on a desktop computer. The human baseline is drawn from a specific pool of workers doing those specific tasks under specific conditions. Crossing that baseline means the model performs better than that reference group on those tasks in those conditions.

It says nothing about whether the model handles the next task it hasn't seen before, applies good judgment in unexpected situations, or reasons about problems that require genuine understanding of context. Those are different questions. No benchmark currently measures them reliably.

Goodhart's Law applies here

When a benchmark becomes a target, it stops being a good measure of the thing it was meant to test. AI labs know which benchmarks matter, and training pipelines can improve scores without improving the underlying capability the benchmark represents. This isn't necessarily intentional: if a model trains on millions of examples resembling benchmark tasks, it gets better at benchmark tasks. Whether it gets better at the real-world equivalent is a separate question that requires separate evaluation.

Independent researchers have repeatedly found that leading models score poorly on benchmark variants they haven't trained on, even when those variants test the same nominal skill. The scores are real. The generalization is not always.

What 'human baseline' actually means

The comparison point matters. Many AI benchmarks use crowd-sourced workers or junior employees as the human reference, not domain experts. A model that outperforms a crowd-sourced worker on a software task is a different achievement than outperforming an experienced software engineer. The phrase "AI beats humans" often describes a more modest result than it sounds.

That's not a reason to dismiss the progress. OSWorld and GDPval scores in the 75-87% range represent genuine capability advances, and AI models are getting better at practical tasks quickly. The problem is the leap from "this model scored higher on this test than these workers" to "AI is now smarter than humans." The evidence doesn't reach that conclusion, and treating benchmarks as if they do tends to produce both breathless coverage and unwarranted alarm rather than accurate understanding of what's actually happening.

The real story is interesting enough

Models that reliably control desktop applications, write production-quality code, and navigate multi-step research tasks are genuinely useful and changing how industries work. That is worth covering carefully. Benchmark scores are one window into that progress, but a narrow one. Reading them as proof that AI thinks the way humans do is a category error the data doesn't support.

Sources

  1. i. openai.com
  2. ii. almcorp.com
  3. iii. hai.stanford.edu

Commentarii · 0

Add · a · Comment