A chatbot can write an explanation even when no webpage answers your exact question. It generates text using patterns learned during training and the information in your conversation. If it has search or document tools, it can also use material those tools retrieve. A conventional search engine mainly finds and ranks existing pages; many search products now combine that with generated answers.
The ability to produce an answer does not mean the answer is correct. Without retrieval, a chatbot usually cannot identify an original source for each claim. It can combine ideas helpfully, but it can also invent a plausible detail. Google’s introduction to LLMs explains how text prediction works; NIST’s Generative AI Profile describes the risk of confidently false output.
Use a chatbot to explain, compare, and organize information. When a factual claim matters, open its sources and check that they support the claim. If it supplies no reliable source, verify the answer elsewhere.
Search retrieves documents; a chatbot writes an answer
A conventional search engine crawls and indexes pages, then ranks results for your query. This is useful when you need an original document, manual, official form, or current announcement. Google’s explanation of search describes these stages.
A chatbot generates a response one token at a time. It can explain an insurance letter, compare phone plans you paste into the conversation, or suggest meals from a list of ingredients. It can do this without finding a page that matches your request exactly. Google’s explanation of LLMs covers the underlying process.
The two approaches often work together. For a delayed train journey, use search to find the operator’s current compensation rules. A chatbot can then help explain those rules or draft a claim using your journey details. Check the operator, ticket conditions, dates, and applicable rules before submitting it.
“Sources” can mean five different things
People use “Where did the AI get that?” to mean several different things. Separating them prevents a common misunderstanding.
1. Training data: material used to teach the model patterns
Before release, a language model is trained on very large collections of text and often other data. During training, it adjusts millions or billions of numbers, called parameters, to make better predictions. Further training can teach it to follow instructions or behave better on particular tasks. Google's explanation of training and fine-tuning, Google's fine-tuning guide.
A response usually cannot be traced sentence by sentence to its training data. The model does not normally look up an original training document while answering, nor can it reliably say “this sentence came from page X.” Knowledge from training is distributed across parameter values rather than kept as a neat card catalogue. The research paper that introduced retrieval-augmented generation explicitly distinguishes a model's parameter-based memory from an external document index, and identifies provenance as a difficulty for parameter-only knowledge. Lewis et al., Retrieval-Augmented Generation.
So, “my training data includes the source” and “this answer is supported by that source today” are very different claims. Treat a base-model answer as a useful draft or explanation, not as evidence that you can cite.
2. Parameters: the model's learned, compressed patterns
Parameters are the adjustable numbers that training changes. A helpful but imperfect analogy is that training turns many examples into learned habits, not into a searchable shelf of books. The parameters can help the model explain photosynthesis, write a summary, or recognise a familiar pattern. They do not provide a dependable list of the documents that support an individual sentence.
This also explains a surprising behaviour: the model may state a correct fact in different words, combine facts it saw in different contexts, or make up a plausible-sounding detail when the pattern is weak. It is generating a continuation, not quoting a verified database record.
3. Your current context: what you supply in the conversation
The prompt, earlier messages, pasted text, uploaded files, and sometimes connected workplace data can be direct inputs to an answer. This is the most traceable kind of source if you can inspect the material yourself.
For example, if you paste three electricity bills and ask for a comparison, the chatbot may calculate from those bills. It may still make a transcription or arithmetic error, so check the numbers, but its evidential basis is clearer than a vague answer from training.
Do not paste private data merely because a chatbot is convenient. Confirm the product's data and workspace settings before sharing personal, medical, financial, client, or confidential information.
4. Retrieved documents: web search, a knowledge base, or a file search
Some chatbots can search the live web, search an index, or retrieve passages from files before drafting. This pattern is often called retrieval-augmented generation, or RAG. The retrieval system finds possibly relevant passages; the language model uses those passages and its general language ability to write the response. The classic RAG work combines a parameter-based model with an external document index. Lewis et al..
This can make an answer more current and auditable, but it is not a truth guarantee. The system may retrieve poor sources, miss the best source, misunderstand a passage, mix supported and unsupported claims, or use a source that has changed. An answer can have ten citations and still fail if none supports its key conclusion.
As one concrete product example, OpenAI says that ChatGPT's web-search answers may include citations and that readers should open them, check publication or update dates, and prefer authoritative sources when accuracy matters. OpenAI: Searching the web with ChatGPT. Other products have different retrieval systems and source displays, so check the specific product's documentation instead of assuming their behaviour is identical.
5. Citations: links intended to let you check claims
A citation is a claimed connection between an answer and a source. It is the most useful form of “source” for a reader, provided it is real and correctly placed. It does not prove that every surrounding sentence is supported. It may point to a page that is irrelevant, outdated, paywalled, misquoted, or lower quality than a primary source.
Open the citation and check the passage that supports the claim. NIST warns that generative systems can produce confidently false content, including false citations or reasoning that appears to justify an incorrect answer. NIST Generative AI Profile, section 2.2.
Why a smooth answer can be wrong
“Hallucination” is the popular term for an answer that is invented or false but sounds plausible. NIST uses the more neutral word confabulation for confidently stated erroneous or false content. NIST Generative AI Profile. The failure is not necessarily a deliberate deception. The model is optimising for a likely, helpful continuation of the conversation, not independently proving each proposition.
Common causes include:
- The question assumes a false premise, but the model accepts it instead of challenging it.
- The fact has changed since training, such as a price, office-holder, policy, product feature, or timetable.
- The prompt lacks location, date, product version, or other decisive context.
- The answer requires an exact quotation, calculation, legal rule, or citation that the model has not retrieved.
- Retrieved sources disagree, are low quality, or are being summarised too broadly.
Fluency is therefore a poor reliability signal. The better signals are a specific claim, a relevant primary source, a visible date and jurisdiction where relevant, agreement with another independent authoritative source, and a calculation you can reproduce.
Base-model answers and web-enabled answers are different modes
Before relying on a response, establish which mode you are in.
| Mode | What is likely informing the response | Strength | Main limitation | Best use |
|---|---|---|---|---|
| Base model, no retrieval | Training-derived patterns plus your conversation | Fast explanation, drafting, brainstorming, translation, reformatting. | May be stale; no inspectable evidence trail for a factual claim. | Low-stakes work and first drafts. |
| Model with pasted text or files | The supplied material plus the model | Can organise and explain a defined set of documents. | May omit, misread, or invent beyond the material. | Summaries and comparisons, with spot checks. |
| Model with web or knowledge-base retrieval | Retrieved pages or passages, the current context, and the model | Can combine current material and provide links. | Retrieval, source quality, and synthesis can all fail. | Research assistance, provided citations are checked. |
The modes can overlap. A web-enabled chatbot does not stop using its learned parameters when it searches, and a base model may have learned some old version of a fact. The practical rule is simple: if a claim needs to be current, attributable, or defensible, ask for sources and verify the important ones directly.
A practical workflow for checking an AI answer
Use this workflow when an answer will influence a purchase, travel, health decision, legal question, school work, work output, or public claim.
- Classify the claim. Is it a creative suggestion, a stable explanation, a current fact, a calculation, or a high-stakes instruction? The more current or consequential it is, the more checking it deserves.
- Ask for the evidence you actually need. Try: “Use current primary sources. Give a link directly after each factual claim. State the country, date, and any assumptions.” For a calculation, ask it to show inputs and arithmetic.
- Open the two or three most important citations. Start with an official regulator, manufacturer, public body, original study, or contract, not a blog that repeats one.
- Check that the source supports the claim. Find the sentence, table, or rule in the source. Does it really say what the chatbot says? A source about delays in general does not prove a refund for your ticket.
- Check scope and freshness. Note the publication or update date, country, product version, eligibility criteria, and exceptions. Search ranking and a citation list do not replace this contextual check.
- Cross-check a consequential conclusion. Use a second independent authoritative source, or contact the responsible organisation. For medical, legal, tax, safety, or investment decisions, professional advice or the official authority may be necessary.
- Keep the source, not only the summary. Bookmark or save the official page and record its date if you need to defend the decision later.
If the chatbot gives no sources
Do not ask only “Are you sure?” That often produces more fluent wording rather than better evidence. Instead:
- Ask whether the answer was based on web search, supplied files, connected data, or general model knowledge.
- Ask it to separate claims that need current evidence from its reasoning or recommendations.
- Request direct, primary-source URLs, dates, and quotations no longer than necessary to locate the relevant passage.
- If it cannot provide checkable evidence, treat its output as an unverified explanation. Search for the primary source yourself or use a search-enabled research tool.
For a low-stakes task such as a meal plan, that may be enough. For a decision with legal, medical, financial, privacy, or safety consequences, lack of sources is a reason to pause, not a reason to trust confidence.
If a citation does not support the claim
Open the link. Check whether it is the correct page, whether it is current, and whether it supports the exact claim rather than a nearby but different one. Then tell the chatbot precisely what failed:
“This link describes general passenger rights, but it does not state that my ticket is refundable. Revise the answer using the operator's current conditions of carriage and quote the relevant clause.”
Do not accept a replacement citation without opening that one too. A source that does not support a claim should be removed or the claim narrowed. If the sources conflict, say so in the final wording rather than forcing a single confident conclusion.
A sensible division of labour
Choose the tool according to what you need from it.
| Your task | Search can help you | A chatbot can help you |
|---|---|---|
| Understand an unfamiliar concept | Find original explanations and alternative perspectives | Work through an explanation or example |
| Check a current rule | Locate the official rule and its effective date | Explain supplied wording and draft questions |
| Use material you already have | Find additional evidence to compare | Organize it into a draft, checklist, or comparison |
Keep the underlying sources visible. Read important passages yourself, and use the relevant professional or official authority when a personal decision requires it.
Evidence
Sources used for this answer.
Question signals show what people need. Primary documentation supports the answer. Both remain visible.
- 01Perché l'ia mi sa rispondere mentre il motore di ricerca no? Quali sono le sue fonti?Reddit · question signal · checked 25 Aug 2026
- 02Google’s explanation of searchdevelopers.google.com · implementation guidance · checked 25 Aug 2026
- 03Google Search Central: A guide to Google Search ranking systemsdevelopers.google.com · implementation guidance · checked 25 Aug 2026
- 04Vaswani et al.: Attention Is All You Needarxiv.org · primary evidence · checked 25 Aug 2026
- 05Google’s introduction to LLMsdevelopers.google.com · implementation guidance · checked 25 Aug 2026
- 06Google's fine-tuning guidedevelopers.google.com · implementation guidance · checked 25 Aug 2026
- 07Lewis et al., Retrieval-Augmented Generationarxiv.org · primary evidence · checked 25 Aug 2026
- 08OpenAI: Searching the web with ChatGPThelp.openai.com · implementation guidance · checked 25 Aug 2026
- 09NIST’s Generative AI Profiletsapps.nist.gov · primary evidence · checked 25 Aug 2026