AI is genuinely excellent at research and genuinely dangerous as a source, and those two facts are not in tension. They describe different jobs.
Almost everyone who gets burned has quietly merged them.
The distinction the whole thing rests on
Two questions that sound similar and are not:
"Help me find it." Point me at the thing. Give me the search terms I did not know. Tell me who works on this. Summarize what I already have in front of me.
"Tell me the answer." What is the number. When did it happen. What does the ruling say. Is this legal.
The first is a navigation problem, and models are superb at it. They have read enormously more than you and can name the field, the term of art, the obvious prior work, the thing you should have searched instead. Being slightly wrong costs nothing, because you are about to go look.
The second is a retrieval problem, and it is the one that burns people. A model producing a citation is not looking anything up. It is producing what a citation looks like, which is a formatted string of plausible-looking parts. Sometimes those parts describe a real source. Sometimes they do not, and nothing about the output distinguishes the two cases.
Search-connected tools change this genuinely but not completely: now some of the claims are grounded in a retrieved page and some are still the model talking, and they arrive in the same paragraph, in the same voice.
The workflow
Four steps. The whole thing is arranged so that no claim reaches your work without a human having seen its source.
1. Ask it to map, not to answer. Open with orientation rather than conclusion:
I need to understand how [X] works. Do not tell me the answer yet.
Give me the terminology I should be searching, the main positions
people take, and what a person new to this usually gets wrong.
You now have search terms, not facts. Nothing here can burn you because nothing here is load-bearing yet.
2. Get the sources, then read them. Ask for what to read, then go and read it. The model's summary of a paper is a hypothesis about the paper. It is a good hypothesis. It is not the paper.
3. Put the material in the window. Once you have the actual documents, paste them in and ask your questions against them. This is the mode where models are strongest and safest, because you have moved the job from recall to reading comprehension. Context is everything applies here more than anywhere.
4. Ask what would change the conclusion. Before you use it:
What would have to be true for this to be wrong? What is the strongest
argument against it, and what evidence would settle it?
Making citations checkable
If you are going to ask for sources anyway, ask in a shape that makes verifying them cheap. Three things do most of the work:
Demand identifiers, not prose. A DOI, an arXiv number, a case citation, a section number. Free-form titles and author lists are the easiest thing in the world to generate plausibly. An identifier either resolves or it does not, and checking takes seconds.
Ask for the quote. Requesting the exact sentence that supports the claim is a strong filter, because a fabricated source rarely survives being asked to produce specific supporting text, and when it does the fabrication is usually visible.
Make it mark its own confidence. Explicitly:
For each claim, mark it [SOURCED] if it comes from a document I gave
you, [RECALLED] if it is from training, or [INFERRED] if you worked
it out. Do not blur the three.
This is not a guarantee. The labels are generated like everything else. But it reliably surfaces the model's own uncertainty, which is otherwise flattened into one confident register, and the RECALLED lines are your checking list.
Claims to never accept without opening the source
These are the categories where the failure is both most likely and most expensive. Every one is specific, checkable, and rewarding to fabricate.
- Numbers with decimal places. Precision reads as authority. "Roughly a third" is a summary; "34.7%" is a claim about a specific table in a specific document.
- Anything legal. Case names, statutes, section numbers. The Mata v. Avianca sanctions were about exactly this, and the lawyers were not careless people, they were people who did not know this failure mode existed.
- Medical and dosage information. No qualifications needed here beyond the cost of being wrong.
- Quotes attributed to a named person. Plausible-sounding quotes are among the easiest things to generate and among the most damaging to publish.
- Anything after the training cutoff. Recent events are a structural blind spot, and the model may not flag it.
- Any claim you want to be true. Not a property of the model, a property of you. The claim that supports your argument is the one you will check least.
What this actually costs
Roughly nothing, which is the point people miss. Steps 1 and 3 are where models are most useful and least risky, and they are the majority of research time. The discipline is concentrated in a small number of load-bearing claims.
You are not adding verification to everything. You are noticing which handful of sentences are doing the work, and checking those.
The short version
Use it to find things, not to know things. Get identifiers rather than titles. Read the source rather than the summary of it. Once you have the documents, put them in the window and work from there. Check numbers, law, medicine, quotes, recent events, and anything you were hoping was true. Verify by opening the source, never by asking again.