The disease is called bixonimania. The symptoms, as the academic literature would have it, include eyelid discoloration and sore eyes brought on by prolonged exposure to blue light from mobile devices. It is not a condition recognised by any medical body, because it does not exist. A researcher invented it in 2024 to find out how easily large language models would absorb a piece of fabricated medical information and start repeating it to users as fact.
The answer, set out in a recent investigation by Nature, is that they absorbed it almost immediately.
The experiment
Almira Osmanovic Thunström, a researcher at the University of Gothenburg, uploaded two fake academic papers describing bixonimania to a preprint server. Preprint servers do not peer-review submissions, but their content is routinely indexed and pulled into the training pipelines and retrieval systems used by major chatbots. Within weeks, ChatGPT, Google Gemini, Microsoft Copilot and Perplexity were citing bixonimania as a real condition, sometimes confidently describing its causes, symptoms and recommended treatments.
The contamination did not stop with the chatbots. A subsequent paper in Cureus, a journal owned by Springer Nature, cited bixonimania in its own discussion of digital-age eye conditions. After Nature contacted the journal for comment, the paper was retracted on 30 March 2026. Coverage of the fallout is available at Nurse.Org and Futurism.
The finding underneath the prank
This story keeps getting trimmed down to "AI fell for a fake disease," which misses the structural finding underneath. The chatbots were notably more likely to hallucinate around bixonimania when the prompt was framed in formal clinical language. Nature reports that when the input looked like a hospital discharge note or a clinical paper, the hallucination rate rose.
In other words, the register that would make a human reader more cautious, the register of a clinical paper or hospital note, makes a chatbot less cautious. That inverts the trust signal patients are usually told to look for. The model treats authoritative-sounding language as evidence that the underlying content is also authoritative, which is the opposite of how a clinician is trained to read it.
What the case does and doesn't prove
It is worth being careful about the claim. Bixonimania does not show that AI chatbots are useless for medical questions, or that they always invent diagnoses. It shows something narrower and more uncomfortable: that the current generation of frontier models will fabricate clinical detail with the same confidence whether the underlying condition is real or invented, and that the contamination can leak from the chatbot back into the peer-reviewed literature it draws from.
The broader myth worth examining is the still-common claim that today's chatbots are reliable enough to function as a first-line reference for symptom checking. That claim is not settled by a single experiment. But the bixonimania case does undercut the related assumption that, as models grow larger, they grow more grounded. They mainly grow more fluent. Compare with the equally persistent myth that today's chatbots are conscious: in both cases, the fluency of the output is doing most of the persuading.
The practical takeaway is unglamorous. If a chatbot offers medical information in confident, clinical prose, that is not yet a reason to trust it. It may be a reason to check the source it claims to be citing, on the off chance the source was a paper a researcher uploaded as a stress test.
Sources
- i. www.nature.com
- ii. en.wikipedia.org
- iii. theheartysoul.com
- iv. nurse.org
- v. futurism.com
- vi. theprint.in
- vii. www.linkcentre.com
Commentarii · 0