Friday, September 25, 2026

Why AI Hallucinates

A 30-second Short can show the gap being filled. It cannot show you how to catch the fill. If you already watched the video, this post is the part that does not fit in a caption: the mechanism, the incentives, a case that actually reached a courtroom, and a checklist you can use the next time an answer sounds too sure.

What's actually happening

A language model is not looking up a fact and then deciding whether to speak. In pretraining it is asked, over and over, to guess the next token in a stretch of text. Grammar, spelling, and common phrasing are easy because they repeat. A one-off detail — a birthday, a case caption, a paper title that appeared once — is not a pattern. When the prompt still demands a completion, the model fills the hole with the closest fluent continuation it has.

That is gap-filling, not lying. There is no inner judge that knows the difference between "this is attested" and "this would sound right." The same machinery that writes a grammatical sentence will invent a plausible citation. The Short's picture is accurate: a small blank, a huge teal network, a confident plug. The blog's extra space is this: the plug is a statistical best guess, not a retrieval from a verified store.

Ask for a name, a number, a holding, or a URL and you are asking it to complete a rare fact. If that fact was weakly represented in training, the completion will still look like the rest of the sentence — calm, specific, and wrong.

What the research says

OpenAI, 5 September 2025 (Why language models hallucinate): hallucinations are plausible but false statements. The authors argue they persist because standard training and evaluation reward guessing over acknowledging uncertainty. Graded only on accuracy, a model that guesses "September 10" for an unknown birthday has a 1-in-365 shot at a point; "I don't know" is a guaranteed zero. Over thousands of items, the guessing model wins the scoreboard.

https://openai.com/index/why-language-models-hallucinate/

The same team (Kalai, Nachum, Vempala, Zhang) posted the paper on 4 September 2025 as arXiv 2509.04664. If a system cannot tell a valid statement from an invalid one, next-word training produces errors the way a classifier produces false positives. Arbitrary low-frequency facts cannot be predicted from pattern alone. Their proposed fix is socio-technical: change how dominant benchmarks score abstention, not only add another hallucination test.

https://arxiv.org/abs/2509.04664

The cost of a fluent gap-fill is not theoretical. In Mata v. Avianca (S.D.N.Y., 2023), a personal-injury brief cited more than half a dozen decisions — Varghese v. China Southern Airlines, Martinez v. Delta Air Lines, and others — that opposing counsel and the judge could not find. ChatGPT had invented them, then insisted they were in Westlaw and LexisNexis. Judge P. Kevin Castel described bogus opinions with bogus quotes. The lawyers were sanctioned.

https://www.nytimes.com/2023/05/27/nyregion/avianca-airline-lawsuit-chatgpt.html

A confident citation is exactly the plug. The courtroom is what happens when nobody checks the hole.

Why it can't just be "fixed"

People treat hallucination as a bug to patch: more data, a bigger model, a retrieval plugin, a "don't make things up" system prompt. Those help at the margin. They do not remove the incentive.

Pretraining never labels each sentence true or false. It only asks for fluent continuation. Then evaluation, the thing labs actually optimize for, mostly scores exact-match accuracy. A model that abstains looks worse than a model that guesses, even when the guesser is wrong more often in the cases that matter. Leaderboards that give zero for silence and one for a lucky hit will keep producing systems that sound sure.

You can fine-tune a model to say "I'm not sure" on some prompts. You cannot make "I'm not sure" win an accuracy-only exam. Until the tests that define "better" stop punishing honesty, the product that ships will still fill the gap. OpenAI's paper is explicit: this is not solved by one more hallucination benchmark on top of the old scoreboards.

How to spot and handle it

Treat fluent specifics as unpaid invoices. Before you reuse an AI answer:

  • Names, dates, dollar amounts, case captions, paper titles, and URLs get checked in a primary source — a docket, a journal page, a company filing — not by asking the same chat "are you sure?"
  • If the model cites a case or a paper, search the identifier yourself. Mata is what happens when verification is another prompt to the inventor.
  • Ask for the source, then open the source. A quote that cannot be found in the linked page is a fill.
  • For anything you would sign, file, or publish, assume the gap was filled until a human-readable record says otherwise.

None of that makes the model honest. It makes you the person who still owns the merge.

Closing

The next generation of models will be smoother. The scoreboards that train them may still pay for a guess. The useful question is not whether the gap will keep getting filled — it will — but whether your workflow still has a human who is allowed to leave the answer blank.

No comments:

Post a Comment

Why AI Hallucinates

A 30-second Short can show the gap being filled. It cannot show you how to catch the fill. If you already watched the video, this post is th...