AI Literature Search: Why Plausible Citations Are Not Evidence Coverage
Anas H. Alzahrani, MD PhD MPH
Department of Preventive Medicine and Public Health
Faculty of Medicine, King Abdulaziz University
An AI literature search can return ten real clinical studies, summarize them accurately, and still miss the evidence that would change the conclusion. Valid citations answer “Are these papers real?” Evidence coverage asks the harder question: “Did we find enough of the eligible studies to trust the synthesis?”
Clinical researchers should treat fluent citation retrieval as a search layer, not a completed review. The methodological risk is no longer limited to fabricated references. It includes invisible omissions, unstable retrieval, eligibility errors, and a search process nobody can reproduce.
The Clean Metaphor: A Bookshelf Is Not an Inventory
A shelf of genuine books can still be a poor inventory of the library.
Checking each title proves that the books exist. It does not reveal the empty shelves. AI answers are easy to judge by what they show and difficult to judge by the eligible evidence they never retrieved.
Four Questions That Should Never Collapse into One
| Question | What it tests | What it cannot prove |
|---|---|---|
| Is the citation real? | Bibliographic validity | Eligibility or coverage |
| Is it relevant? | Topical fit | Protocol eligibility |
| Was it eligible? | Population, intervention, comparator, outcome, and design criteria | Whether other eligible studies were missed |
| How much was recovered? | Recall against a reference set | That the reference set is complete |
A chatbot may perform well on the first question and poorly on the fourth. That distinction matters because omitted small, null, older, or less visible studies can alter effect size, heterogeneity, harms, and certainty even when every displayed citation is authentic.
Interactive retrieval audit
Count What the Answer Missed
Compare a citation list with a known eligible reference set. These synthetic scenarios separate citation plausibility from evidence coverage. The thresholds below are teaching cues, not universal review standards.
The answer cites real, relevant-looking trials but retrieves fewer than half of the eligible benchmark set.
Reference-set recall
40%
12 eligible studies missing
Methods reading
Not defensible as a standalone evidence search. A fluent answer is masking substantial retrieval loss.
The eligibility rate is calculated only among citations matched to the benchmark. It is not full search precision when other citations remain unclassified.
What a 2026 Benchmark Actually Found
A recent preprint tested three general-purpose chatbots on clinical questions adapted from 20 Cochrane reviews. Across models, user roles, and repeated prompts, a response retrieved an average of 39.2% of the studies included by the expert reviews. The systems also cited some studies that the reviews had excluded. Retrieval differed by model and by the role stated in the prompt.
After adjustment for publication year, citations per year, and open-access status, larger sample size was the only independently significant study-level predictor of retrieval. That is a visibility warning: the studies easiest for a model to surface are not necessarily a representative sample of the eligible evidence.
Read the benchmark conservatively
It is a preprint evaluating particular models, versions, prompts, questions, and a Cochrane-derived reference standard at one point in time. The exact percentages will move. The audit principle will not: retrieval needs a denominator and a reproducible trail.
Why One-Shot Answers Lose Studies
A systematic search deliberately favors sensitivity. It expands synonyms, combines controlled vocabulary with free text, searches more than one appropriate source, checks registers and citations when needed, and records the path from records to included studies. A chatbot answer compresses query formulation, retrieval, ranking, eligibility judgment, and synthesis into one opaque interaction.
That compression feels efficient because the rejected and unretrieved records disappear. It is also why repeating the prompt is not a substitute for a search strategy. More runs may add citations, but without deduplication, eligibility adjudication, and a reference denominator, the researcher still cannot tell whether the evidence base is becoming complete or merely longer.
Decision Rules for Clinical Researchers
- For orientation, use AI freely but label the output correctly. It can generate concepts, candidate synonyms, seed papers, and questions for an information specialist. Call that scoping, not comprehensive retrieval.
- For a systematic claim, preserve a reproducible search. Record every source and platform, full search strategy, limits, dates, deduplication method, and supplementary search. PRISMA-S provides a practical reporting floor.
- Separate retrieval from screening. Store records before deciding eligibility. Every exclusion should be traceable to the protocol rather than silently absorbed by the model.
- Benchmark the AI-assisted layer. Test whether it recovers known eligible sentinel studies. Report recall where a defensible reference set exists, and state the limits of that denominator.
- Look for visibility bias. Compare recovered and missed studies by sample size, year, geography, access status, intervention, outcome, and result direction. Coverage can be unequal even when average recall looks acceptable.
- Keep humans accountable for the final evidence set. AI can support search translation, prioritization, or screening, but the review team still owns eligibility, completeness, and the clinical conclusion.
Reviewer Red-Flag Checklist
One red flag does not prove the synthesis is wrong. Several together mean the evidence set is not auditable enough to support a claim of comprehensiveness.
Why This Matters for Aqrab
Evidence coverage is a document-chain problem. The protocol defines eligibility, the search log defines what could be found, the screening record explains what was excluded, and the manuscript turns the surviving studies into a claim. A rigorous critique reconnects that chain instead of grading the prose.
Use Aqrab Try to pressure-test whether an AI-assisted review supports its claims with a reproducible retrieval and eligibility trail. The key reviewer question is simple: what evidence could be missing, and how would we know?
Methods Sources
- Liu Q, Jin Q, Menke JD, Kahnt T, Lu Z. Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions. arXiv preprint. 2026.
- Lefebvre C, et al. Searching for and selecting studies. In: Cochrane Handbook for Systematic Reviews of Interventions, version 6.5.1. 2025.
- Rethlefsen ML, et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Systematic Reviews. 2021;10:39.
- Page MJ, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. 2021.
Keep reading
Don't stop at one method.
Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.
Estimands in Meta-Analysis: When a Shared PICO Still Pools Different Questions
A practical guide to estimands in meta-analysis. Learn how treatment-policy and hypothetical strategies can make trials with the same PICO answer different questions, and how to audit the pool before combining effects.
AI Before–After Studies: When Faster Care Is Not Yet an AI Effect
A practical guide to evaluating healthcare AI after deployment. Audit pre-trends, concurrent comparisons, co-interventions, outcome measurement, and the claim ceiling of before–after and interrupted time-series designs.
AI Surveillance Models: Why a High AUC Cannot Justify Fewer Follow-Up Visits
A practical guide to evaluating AI surveillance models. Learn why high AUC is not enough to reduce follow-up, and audit calibration, thresholds, missed failures, utility, and prospective impact.