You have heard the line, usually delivered with a shrug: nobody knows how these things work, not even the people who built them. It is one of the most repeated claims about modern AI, and it is comforting in a grim sort of way. It also flattens a much more interesting truth. Large language models are hard to understand, but they are not sealed boxes that defy all inspection.
Where the myth comes from
The claim has a real root. A model like the ones behind Claude or ChatGPT is not programmed with rules a human wrote line by line. It learns billions of numerical weights during training, and those weights do not come with labels. Feed in a prompt, get a response, and there is no obvious readout explaining why the model chose those words. That opacity is genuine, and it is why researchers reached for the phrase black box in the first place.
What researchers can actually see
Here is the part the shrug leaves out. Every weight and every internal activation in these models is fully visible to the people running them. The problem was never access. It was interpretation. And on that front the field of mechanistic interpretability has made real progress. Anthropic's team has managed to map millions of concepts inside a production model, identifying internal features that fire for specific ideas and then turning those features up or down to change behaviour. Google DeepMind and a growing community of independent researchers are building similar tools. MIT Technology Review put mechanistic interpretability on its list of breakthrough technologies for 2026, and Anthropic has started using it in practice, checking a model's internal features for deceptive tendencies before release rather than judging it on outputs alone.
The honest verdict
So the myth fails, but so does the opposite comfort. We cannot yet trace most of what a model does on a given prompt, and the tools that work in a research lab do not scale to explaining every answer in real time. Anthropic's own Dario Amodei has called the gap urgent, warning that our understanding lags well behind the capability we keep shipping. It is worth remembering that a model's stated reasoning is not always its real reasoning either.
The accurate picture is neither a black box nor a glass one. It is a dim room where researchers have started switching on lights, one corner at a time. That is a long way from knowing nothing, and a long way from knowing enough.
Sources
- i. www.anthropic.com
- ii. darioamodei.com
- iii. aiweekly.co
Commentarii · 0