Skip to content

Applied AI

Why AI makes things up (and how to work with it)

Why an AI model confidently writes things that are false, where that barely matters, where it is unacceptable, and how to work with the limitation.

5 MIN READ

Anyone who has used an AI chat for a while has run into it: an answer written with complete confidence that turns out to be false. A citation that does not exist, a feature the tool does not have, a regulation that says the opposite of what the model claims. This is not a bug they will fix next month: it is a direct consequence of how these systems work. Which is why the decision that matters is not "can I trust AI?" — in general, no — but in which tasks you can use what it produces without checking every line, and in which you need a different way of working.

It does not look things up: it writes

A language model does not fetch answers from a database of facts. It generates text: it continues what you write with whatever is most plausible given the patterns of everything it has read. Most of the time, the most plausible thing happens to be true — which is why the tool seems to know so much.

But when it has nothing to draw on, the mechanism does not stop: it keeps writing and fills the gap with something that fits. That is where the invented reference with a perfectly credible author, title and year comes from, or the clause it "remembers" from a contract it has never seen. The model is not lying — lying requires knowing the truth. It is doing exactly what it does when it gets things right: producing the most probable text.

Why it sounds so convincing

With people, we use confidence as a signal: those who are unsure, hesitate. A language model breaks that signal. It writes the invented with the same ease, the same professional tone and the same flawless grammar as the correct, because fluency is what it does — not an indicator that anything has been checked. The confidence of the text tells you nothing about the reliability of the content.

There is an aggravating factor too: the more specific your question is about something the model knows little about — your industry, your contract, a recent regulatory change — the bigger the gap it has to fill and the more likely the invention. The very questions that matter most to you are the ones it answers worst from memory.

Where it barely matters

There are tasks where the value lies in the form, not the facts: preparing a draft, rewording a dense text, structuring some notes, suggesting angles on a problem. If the output is a starting point someone will rewrite anyway, an invention gets corrected along the way.

The risk drops especially when you supply the material: "summarise this text", "draft a reply from these notes". The source is right there and checking costs little. It does not vanish entirely — a summary can also add embellishments of its own — so reading before forwarding is still part of the job. But it is a small risk, and an easy one to manage.

Where it is unacceptable

At the other end are the tasks where a plausible invention does real damage: a figure that ends up in a report, a legal or tax reference, the specifications of a product, a fact about a customer, a public reply in the company's name. The test for separating one kind of task from the other has two parts: how much an error costs, and how hard it is to spot. An odd draft stands out on first read; a plausible but false figure inside a report can travel for months without anyone questioning it.

If a task combines high cost and hard detection, the model's answer "from memory" does not count as a source. It can serve as a starting point — never as a claim.

How to work with it

  • Give it the material. The difference between asking "from memory" and asking it to work on a text you provide is enormous. Whenever the document exists, attach it or paste it: the task shifts from "recalling" to "reading".
  • Verify in proportion to the risk. You do not need to check every line of an internal draft; you do need to check every claim in anything that leaves the company or feeds a decision.
  • Ask for what you can check. If a claim matters, ask where it comes from and verify it against a source you control. If there is no way to check it, treat it as a hypothesis, not a fact.
  • Fence the ground when the one answering is not you. An assistant that talks to customers cannot improvise. It is built to answer only from content the company has approved, and to say "I don't know" when the answer is not there. That constraint is not a technical limitation to be papered over: it is the difference between an assistant and a risk with a friendly interface.
  • The signature is not delegated. Whoever sends it, reviews it. AI removes the blank page; responsibility for what is claimed stays with the person claiming it.

Three questions before using an AI answer

  1. Did I give it the material, or did it pull this from memory?
  2. Can I check this against a source I control?
  3. What happens if it is false and nobody catches it in time?

With those three answered, the tool stops being a roulette wheel and becomes what it is: an extremely fast writer that cannot tell the true from the plausible, working for someone who can.

At Dateliers this tends to be the first conversation in any AI training we run: not which buttons the tool has, but what you can and cannot ask of it.

After reading

Does this sound like your case?

If this describes something sitting on your desk, tell us about it. We'll come back with a first read before proposing anything.

Tell us about your case

A first 30-minute call with direct senior interlocution — no commitment and no sales pitch.