Workflows · Intermediate
How to Use AI for Research (A Workflow With Verification Built In)
A six-stage research workflow using AI where it is reliable and avoiding it where it is not — including why you should never ask an AI assistant for citations.
Research is the task where AI assistants are simultaneously most tempting and most dangerous. Tempting because literature work is slow. Dangerous because the single thing these tools fabricate most reliably is a citation.
So this workflow is built around one rule:
Never ask an AI assistant to find sources. Use it to understand sources you found yourself.
That inversion is the whole guide. Everything below follows from it.
Why not ask for sources?
Because fabricated references are indistinguishable from real ones by inspection. A model generates citations by predicting plausible tokens, and citation format is extremely predictable — correct journal names, sensible volume numbers, real researchers who work in the field, page ranges that look right. Everything is correct except existence.
This is not an occasional glitch. It is the expected output of a system that predicts plausible text and has no lookup step. Lawyers have been sanctioned for filing briefs citing invented cases; students have submitted bibliographies where a third of the entries do not exist. In both situations the output looked completely professional.
Use a real index instead — Google Scholar, PubMed, JSTOR, arXiv, your library's database. Those return things that exist. Then bring what you found to the AI.
Stage 1: map the territory (AI is useful here)
Before searching, you often do not know the vocabulary of a field — and search only works if you know the terms.
I'm researching whether four-day work weeks affect productivity in
knowledge work.
I'm not asking for sources. Instead:
1. What academic fields study this, and what does each call it?
2. What are the standard technical terms and their synonyms?
3. What are the main competing positions in this debate?
4. What confounds make this hard to measure?
5. What search terms would a specialist use that a novice wouldn't?
Why it works. This is a vocabulary and structure question, not a factual-recall question. The terminology of a research area is densely represented in training data, so the model is reliable at it. Question 5 is the highest-value one: it converts naive searches into specialist searches, which is the difference between finding review articles and finding nothing.
Question 4 is worth keeping — knowing the confounds before you read means you evaluate studies rather than absorb them.
Stage 2: search real databases (no AI)
Take the terms from stage 1 to actual indexes. Prefer:
- Review articles and meta-analyses first. They map the field and their reference lists are a curated bibliography.
- Citation chaining. Follow references backwards from a key paper; use "cited by" to move forwards.
- The primary source, always. Never cite a paper you have only read described.
This stage is unglamorous and not delegable. What you gain is a set of sources that definitely exist.
Stage 3: triage what to read closely (AI is useful here)
You now have thirty candidate papers and time for six. Paste an abstract:
Here's an abstract: [paste]
My research question: [state it]
1. Is this directly relevant, background, or tangential?
2. What's the study design, and what can it therefore NOT establish?
3. What's the population/sample — does it generalise to my context?
4. Based only on this abstract, what should I be sceptical about?
5. Worth reading in full for my question? One-line answer.
Why it works. The abstract is in front of it, so this is reading comprehension rather than recall — the reliable mode. Question 2 is the one that improves research quality most: a cross-sectional study cannot establish causation, a lab study may not transfer to field conditions, and naming the design's ceiling prevents you from over-claiming later.
Stage 4: interrogate the papers you read (AI is useful here)
For a paper you are reading properly, paste sections and ask:
Here's the methods and results section: [paste]
1. In plain language, what did they actually do?
2. What did they measure, and is it a good proxy for what they claim?
3. What are the stated limitations?
4. What limitations are NOT stated but visible in the method?
5. Do the results support the strength of the conclusion, or is there
a gap between findings and framing?
Quote the specific lines behind each of your points.
Why it works. The "quote the specific lines" requirement is the verification mechanism: if a quote is not in your pasted text, you have caught a fabrication instantly. Question 2 catches proxy problems — measuring self-reported productivity and concluding about actual output. Question 5 catches the gap between what a study found and how its abstract describes it, which is extremely common and hard to notice when you are reading fast.
This stage is genuinely where AI adds the most research value: a tireless critical reader that will go through a methods section with you line by line.
Stage 5: synthesise across sources (AI helps, carefully)
Once you have read several papers and taken your own notes:
Here are my notes on five studies: [paste your notes]
1. Where do they agree?
2. Where do they genuinely conflict — as opposed to measuring
different things?
3. What would explain the conflicts (population, method, period,
definitions)?
4. What does this set NOT address?
5. What's the most defensible summary given only these five?
Use only my notes. Do not add studies or facts I haven't given you.
Why it works. Point 2 is the sharpest distinction in research synthesis — apparent contradictions are usually definitional differences rather than real disagreement, and spotting that is what separates a good literature review from a list. Point 4 identifies your gap, which is often the actual contribution you can make.
The final instruction is a hard boundary and you should include it every time. Without it, models pad synthesis with plausible-sounding studies from memory, and those will be fabricated.
Stage 6: verify before you write
Before anything goes into a document:
- Every citation came from a database you searched, not from an AI
- You have opened every source you cite
- Every quote is checked character-for-character against the original
- Every number is traced to the primary source
- Every claim attributed to a paper is one you personally read there
- Nothing is cited on the basis of an AI summary alone
The fifth item is the one people rationalise past under deadline pressure. It is also the one that produces retractions.
What AI is good and bad at, in one table
| Task | Reliable? | Why |
|---|---|---|
| Explaining terminology and field structure | Yes | Densely represented, low specificity |
| Suggesting search terms | Yes | Vocabulary task |
| Summarising a paper you paste | Yes | Reading comprehension |
| Critiquing a method you paste | Mostly | Pattern recognition against known designs |
| Synthesising your own notes | Mostly | Bounded by material you provide |
| Explaining a statistical method | Mostly | Verify against a textbook |
| Finding sources | No | Fabricates convincing references |
| Providing citations | No | Same mechanism |
| Quoting from memory | No | Reconstructs plausible-sounding quotes |
| Reporting statistics from memory | No | Numbers drift, qualifiers vanish |
| Telling you what current consensus is | No | Training cutoff, no recency awareness |
The pattern is consistent with everything else about these tools: it is reliable when you supply the text and unreliable when it must supply the facts.
Two traps specific to research
The drifting statistic. A model reports "73% of remote workers report higher productivity." The real study said 73% of surveyed respondents at companies over 500 employees reported some improvement in at least one self-assessed area. The number survived; every qualifier that made it meaningful did not. This is undetectable from the output — only opening the source catches it.
The real paper, wrong claim. Subtler than a fabricated citation and harder to catch. The paper exists, the authors are right, the year is right — but it does not say what the summary claims. This is why "verify the citation exists" is insufficient; you must verify it supports the claim.
If web search is involved
When an AI tool searches the web, reliability improves because real text enters the context. But retrieval adds no verification step — the model can still misread a page, merge two sources, or summarise a low-quality one confidently.
For research specifically: web search is useful for orientation and for current information, and it is not a substitute for academic databases. A blog post summarising a study is not the study. Open the cited page every time; if a claim in a search-grounded answer has no citation attached, treat it as unsourced.
Using this with PhantomAI
Stages 1, 3, 4, and 5 are what PhantomAI is actually good for. You can attach papers or paste sections and ask questions against that text — the grounded mode where fabrication risk is lowest. openai/gpt-oss-120b handles the method critique and synthesis work; web search is available when a question needs current information.
To be explicit about what it will not do well: do not ask PhantomAI for citations or reading lists. It will produce convincing, formatted, non-existent references, exactly as described above. That is not a defect specific to this product — it is the mechanism, and it applies to every general-purpose assistant. Use Scholar, PubMed, or your library, then bring the papers here.
See our limitations page for what else the product cannot do, and how to fact-check AI output for the verification methods.
Continue
- How to fact-check AI output — citation, statistic, and quote checks in detail
- Limitations of AI assistants — reproducible failure modes
- AI for studying — retention-focused study workflows
Simanta Pratim Das
Founder & Developer
Simanta is an independent AI engineer based in Guwahati, India, building PhantomAI as a solo project — designing the product, the interface, and the AI pipeline end to end.
Related guides
Try these techniques in PhantomAI
PhantomAI is free to use, and the workflows in this guide work best when you paste in your own material rather than relying on the model's memory. Before you start, it's worth reading what it can't do.