Why it makes things up
A fluent answer is not the same as a true one. Learn why models invent facts, and spot the four places where the risk is highest: citations, exact numbers, niche facts and recent events.
Hallucination
When a model produces information that sounds plausible but is false or invented: a book that doesn't exist, a wrong date, a quote nobody said. It isn't lying, because it has no intent to deceive. It is doing what it always does, producing likely text, in a spot where the likely text happens to be untrue.
You ask for a source and receive an author, a title, a year and a page number. All four look right. The book was never written.
THE CAUSE
Fluent is not the same as true
Remember lesson 1: the model predicts text that fits. It doesn't look facts up in a database and it has no separate alarm that rings when it doesn't know. When the fact is missing or faint in its training, the best-fitting continuation is still a confident, fact-shaped sentence. Good grammar and a sure tone tell you nothing about accuracy.
"The bridge was completed in 18__": a sentence like this wants a year, so a year appears. Whether it's the right year depends on how well that fact was covered in training.
Check yourself
Pouya asks for "a famous quote by a well-known economist about saving". He gets a neat quote with a name attached. He can't find it anywhere online. What most likely happened?
- The quote is real but old enough that it only survives in books, not on pages a search engine has indexed
- It produced a sentence that sounds like that economist, because a quote-shaped answer fit best
- The assistant deliberately tried to mislead him
- His internet search was done wrong
Show the answer
It produced a sentence that sounds like that economist, because a quote-shaped answer fit best
Yes. The prompt asked for a quote, so the most fitting continuation is something that looks like one. Sounding like a person is easy for the model; having actually been said is a separate matter.
Four places where the risk is highest
- Citations and quotes: sources, links, page numbers, who said what. Easy to fake convincingly.
- Exact numbers: statistics, dates, prices, measurements. The model has a feel for the rough size, not a record of the figure.
- Niche facts: a small town, a minor historical figure, a little-known product. Little training text means weak knowledge.
- Recent events: anything near or after the training cutoff from lesson 2.
| You ask for | Risk of invention |
|---|---|
| Rewrite my paragraph more clearly | Low |
| Explain how vaccines work | Low |
| Summarise the text I pasted | Low to medium |
| The population of a small town in 1990 | High |
| Five academic sources with page numbers | Very high |
Check yourself
Negar is writing a report with an assistant's help. Which outputs can she use with a light check, and which must she verify before using?
- A clearer version of her own introduction
- "According to a 2019 study by..."
- A suggested outline with five headings
- "Unemployment in the province was 11.3%"
- A link to a government web page
- Three alternative titles for the report
Show the answer
Light check: A clearer version of her own introduction, A suggested outline with five headings, Three alternative titles for the report
Verify first: "According to a 2019 study by...", "Unemployment in the province was 11.3%", A link to a government web page
Check yourself
If an answer includes very specific details, like an author's full name, a year and a page number, that is a good sign the information is real.
Show the answer
False
Specific details are exactly what a fitting continuation looks like, so invented sources come with them too. Specificity is a matter of style, not evidence. The only proof a source exists is finding the source.
Check yourself
- Prompt: "Explain why the sky is blue." The answer is correct, as it almost always is.
- Prompt: "What year was the old bridge in my home town built, and who designed it?" The answer names a year and an engineer. Both are wrong.
What is the key difference between these two cases?
- The second prompt was too short
- The first topic is written about everywhere, the second almost nowhere, so the model had little to go on
- The second question was asked impolitely
- Models are good at science and bad at history
Show the answer
The first topic is written about everywhere, the second almost nowhere, so the model had little to go on
Exactly. How well a fact was covered in training decides how reliable the model is about it. Common knowledge is solid; local and obscure detail is where invention creeps in.
Check yourself
Elham is preparing slides and wants one striking statistic about water use in farming. The assistant offers "agriculture uses 92% of the country's water". What should she do?
- Find the figure in an official published source and use the number and source from there
- Use it, since it sounds about right
- Ask the assistant "are you sure?" and use it if it stands by the number, since it has no reason to insist on something wrong
- Change it to "about 90%" so it's safer
Show the answer
Find the figure in an official published source and use the number and source from there
Right. The assistant's number is a lead, not a fact. Treat it as a hint about what to search for, then quote the published source.
Lesson recap
- Hallucination is plausible text that happens to be false. It comes from prediction, not from an intent to lie.
- Fluency, confidence and specific detail are style. None of them is evidence.
- Expect trouble with citations and quotes, exact numbers, niche facts and recent events.
- Treat those outputs as leads to check, and never pass on a source you haven't seen yourself.