The headline result is the kind that makes people sit up: a small AI model beating much larger ones. The more interesting result is what it was asked to do. Inherent, a London lab founded by former Google DeepMind researchers, says its agent Faraday outperformed Anthropic's Claude Opus 4.8 and OpenAI's GPT-5.5 at a task that is harder than it sounds, reproducing the findings of published scientific papers without being told the answer in advance.

Reproduction is a good test because it resists memorisation. The agent is handed a paper's setup and has to design and run the work that arrives at the result on its own, choosing sensible experiments and discarding dead ends along the way. Inherent's co-founder and chief scientist Edward Hughes calls the quality it is probing 'research taste', an instinct for which experiments are worth running. 'Many PhD students actually start by doing this,' he told TechCrunch.

What Hughes seemed most keen to stress was the how, not the score. Faraday runs on Qwen 3.6, an open model with 27 billion parameters, a fraction of the size of the frontier systems it was measured against. The team trained it with reinforcement learning aimed specifically at the reproduction task, and leans on OpenAI's coding tools for the software side of the work. 'What was most interesting to us was not so much the result of beating those frontier agents,' he said, 'but actually the way we went about building this.'

Read it with the usual caution

This is a company describing its own system, not an independent benchmark, and the details of the test set matter enormously to how far the claim travels. Reproducing a known published result is also a narrower thing than original discovery, where there is no paper to check your work against. A model with taste for tidy, already-solved problems may still flounder at genuinely open ones.

Even with those caveats, the direction is worth noting. Much of the field's effort has gone into making models bigger. A twelve-person startup getting a compact open model to out-reason systems many times its size, on a task built to reward judgement over recall, is a reminder that architecture and training method still count for a lot. Inherent, which has raised a 50 million dollar seed round, plans to grow to around two dozen people by the end of the year. Whether Faraday holds up under outside scrutiny is the test that matters next.

Sources

  1. i. techcrunch.com

Commentarii · 0

Add · a · Comment