Decide when a model's answer needs checking against the world.
A model's training ended some months ago. You ask who currently holds a particular office. What is it doing?
Yes.It has no clock and no sense that time has passed. The answer is the last state of the world it saw, delivered in the present tense.
Not quite.It is often trained to hedge on this, and the hedge is unreliable — the pattern still points at a name.
Not quite.Only if a search tool is attached. Without one there is nothing to check against.
Put in order what happens when a model is given a search tool.
Tap them in order — first to last.
Question→A query is written→Results pasted into context→Answer from what is in front of it
Retrieval moves the problem to the search. It does not remove it.Yes.The model does not gain knowledge. Something else fetches material and puts it where the model can read it — which is why a bad search produces a confident answer built on bad material.
A tool changes what is in the context window. It never changes what the model is: a predictor working on the text in front of it.
For which of these should you insist on a source before acting?
A dosage or a safety limit.
A rewrite of your own paragraph in plainer words.
A legal deadline for filing something.
A brainstormed list of angles to consider.
A specific figure you are about to publish.
Yes.The line is not how hard the task is. It is whether being wrong costs anything, and whether the answer depends on a specific fact rather than on shape.
The same question, asked in ways that give the model more or less to work with.
Question aloneQuestion + the documentQuestion + document + "quote the line"
8 of 20 · answers that hold up when checked
Question aloneIt answers from patterns. On thin ground, that is where invention lives.
17 of 20 · answers that hold up when checked
Question + the documentThe material is in front of it. Now it is summarising rather than recalling.
19 of 20 · answers that hold up when checked
Question + document + "quote the line"Every claim is anchored to text you can see, so a wrong one is visible immediately.
Why does supplying the document help so much?
Yes.The strongest single habit with these tools: bring the material to the model rather than asking it to remember. It also makes checking cheap, because the source is on your screen.
Not quite.It learns nothing. The document is in the context for this conversation and gone afterwards.
Not quite.Care is not a setting. What changed is what is available to predict from.
Move the control to see what changes.
You need a fact and you have no way to check it. What is the reasonable move?
Yes.It keeps the usefulness — a lead is genuinely valuable — without passing an unchecked prediction on as a fact. The failure mode is not using the tool; it is laundering its output.
Not quite.Confidence is generated by the same process on thin ground as on thick. You know this one now.
Not quite.Three draws from the same distribution agree with each other, not with the world.
Which sentence would you now be willing to defend?
Yes.It explains the fluency, the range, the invention, the sycophancy, the letter-counting and the stale facts — all from one mechanism, which is what a good explanation does.
Not quite.A search engine returns documents that exist. This produces text that is likely, which is a different thing and fails differently.
Not quite.The failures are wrong for that. People do not miscount letters in a word they can see, and they do not invent citations with a straight face.
Lesson complete
Bring the material to the model instead of asking it to remember.