Back to blog
Industries30 templates

30 AI Prompts for Literature Review & Research Synthesis (2026)

30 copy-paste AI prompts for literature review: scoping, summarization, comparison matrices, gap analysis, and methodology critique. Honesty rules included.

30 templates · free to copy · no signup

Showing all 30 templates

TL;DR: Here are 30 AI prompts for research organized across five literature review stages — scoping, summarization, comparison matrices, gap analysis, and methodology critique. Every prompt works inside ChatGPT, Claude, or Gemini. AI hallucination is real: these prompts are designed to work from text you paste in, not from the model's memory. Citation verification stays with you.

What are the best AI prompts for research and literature review in 2026?

The best AI prompts for research are stage-specific templates that give the model a bounded task within your existing workflow — not open-ended instructions to find or generate sources. The distinction matters because AI hallucination is a real risk: models asked to produce citations from memory will generate plausible-sounding references that do not exist. The 30 prompts in this guide work from papers you have already found and imported; the model's job is structure and synthesis, not discovery.

Researchers consistently lose the most time at five stages: deciding what a literature review needs to cover (scoping), extracting consistent information from individual papers (summarization), comparing findings across studies (comparison matrices), identifying what is missing or contradicted (gap analysis), and evaluating the quality of the evidence base (methodology critique). Most AI prompt guides for literature review cover summarization and stop there. This guide covers all five stages, with particular attention to the analytical layer — comparison matrices and gap analysis — that competing resources skip.

Before running any of these prompts: import your papers into Zotero or another reference manager, collect abstracts and key sections, and keep your database search results strictly separate from anything AI-generated. That separation is what makes the output usable in a submitted paper. For connecting these prompts into a versioned, reproducible system that satisfies methods-section disclosure requirements, see our guide on building a reproducible AI research workflow.

What do most AI prompt guides for literature review miss?

Most guides stop at two prompt types: "summarize this paper" and "write an outline for my literature review on [topic]." The second type is the dangerous one — asking for an outline on a topic without pasted evidence is an invitation for the model to invent the field.

The analytical layer is where AI adds the most value for experienced researchers, and it is where guides consistently underdeliver. Building comparison matrices across ten studies used to take a day of manual work. Gap analysis requires holding contradictions across a body of evidence in mind simultaneously. Methodology critique means applying consistent quality criteria (sample size, control conditions, replication, population) across papers that describe their methods differently. AI can apply a consistent framework to a large text set faster than any researcher can — but only if you give it the right framework and the actual text.

The honesty principle running through all 30 prompts: AI structures and synthesizes what you provide; you verify every factual claim against the original source.

StagePromptsAI producesYou verify
Scoping1–6Research question variants, inclusion criteria, search stringsCoverage against field standards
Summarization7–12Per-paper extractions, structured abstractsAccuracy against the original paper
Comparison13–18Cross-study matrices, thematic clustersFair and accurate attribution
Gap analysis19–24Contradiction lists, understudied populationsWhether gaps are real or evidence artifacts
Critique25–30Bias checklists, quality flagsYour judgment on whether issues matter

What prompt structure works best for academic research tasks?

Three components make a research prompt reliable across all five stages. We call this the research prompt layer system.

The first layer is a role instruction that bounds the model to provided materials: "Work only from the text I paste below. Do not add claims from your training knowledge." Without this instruction, the model draws on general training knowledge and mixes it with your provided content — the condition that produces hallucinated facts.

The second layer is the actual content you paste: abstracts, methods sections, structured summaries, or extracted claims. The quality of your paste determines the quality of the output. A full abstract produces a better extraction than a paper title.

The third layer is a specific output format: a comparison table, a bulleted gaps list, a numbered checklist. Specifying the format means the output slots directly into your review document rather than requiring reformatting. Every prompt below follows this three-layer structure.

How do you scope a literature review with AI prompts?

Scoping defines what a review needs to cover and sets the boundaries of the evidence search. Poorly scoped reviews either miss critical evidence or include irrelevant studies. These six prompts sharpen your question and build defensible inclusion criteria before you run a single database search.

1. Research question framing

Use when you have a topic and need it narrowed into something a review can actually answer.

Role: Research methodologist who has supervised systematic reviews and has
seen more of them fail from a vague question than from weak searching.

Context
- Topic: [what you're interested in]
- Field/discipline: [sets the conventions]
- Review type: [systematic / scoping / narrative / rapid]
- Constraints: [time, access to databases, team size]
- What I already know: [prior reading, or "almost nothing yet"]

Task
Turn this into 3 candidate review questions, then frame the strongest one.

Rules
- Frame with the structure the review type expects — PICO for intervention,
  PEO for qualitative, SPIDER for mixed. Name which you used and why.
- Each question must be answerable by reading literature. If answering it
  would require new primary data, say so and rewrite it.
- Do not cite papers, authors, or findings. You have not been given any, and
  a fabricated citation here poisons everything downstream.
- Say plainly if the topic is too broad for the stated constraints, and give
  the narrower version.

Example of the standard I want
Weak:   "What is the impact of social media on mental health?"
Strong: "Among adolescents aged 13-18, is daily active social media use
         associated with higher self-reported anxiety compared with passive
         use, in observational studies since 2015?"

Output
3 candidate questions, each with framework and scope · the strongest one with
reasoning · what it deliberately excludes · roughly how large a literature
that scope implies.

Use this before your first database search. Five variants surface scope problems while you can still fix them without wasted search time.

2. Inclusion and exclusion criteria

Use when you have a question and need screening rules you can apply consistently.

Role: Systematic review methodologist. You write criteria a second screener
could apply without asking you what you meant.

Context
- Review question: [paste it]
- Field: [discipline]
- Review type: [systematic / scoping / rapid]
- Practical limits: [languages you read, date range, access]
- Known edge cases: [study types you're unsure about]

Task
Write inclusion and exclusion criteria.

Rules
- Every criterion must be decidable from a title and abstract alone, or marked
  as a full-text criterion. Screening stalls where that line is blurry.
- Give the reason for each limit. A date cut-off or language restriction that
  isn't justified in the protocol becomes a reviewer's first question.
- Cover the awkward cases explicitly: conference abstracts, preprints,
  dissertations, retracted papers, secondary analyses of the same cohort.
- Flag any criterion that risks systematic bias — excluding non-English work
  is a known source, so name it as a limitation rather than hiding it.
- Do not invent citations or prevalence figures.

Example of the standard I want
Weak:   "Exclude low-quality studies"
Strong: "Exclude studies without a comparison group (full-text stage).
         Rationale: the question is comparative; single-arm designs cannot
         answer it."

Output
Inclusion table and exclusion table, each with rationale and screening stage ·
edge-case decisions · the limitations this introduces.

3. Boolean search string construction

Use when your criteria are settled and you need the query that finds the papers.

Role: Research librarian who builds reproducible searches and reports them so
others can rerun them.

Context
- Review question: [paste it]
- Key concepts: [the 2-4 concept blocks]
- Databases: [PubMed, Scopus, Web of Science, PsycINFO, IEEE, etc.]
- Date range and languages: [and why]
- Papers I know should be found: [2-3 known-relevant studies, if you have any]

Task
Build the Boolean search string for each database.

Rules
- One block per concept, ORs inside a block, ANDs between blocks. Include
  synonyms, spelling variants (behaviour/behavior), acronyms and plurals.
- Adapt syntax per database — truncation symbols, field tags and proximity
  operators differ, and a string copied between them silently fails.
- Mark controlled vocabulary (MeSH, Emtree) as [VERIFY IN DATABASE]. Term
  trees change and you cannot confirm current headings from here.
- Do not claim result counts. You cannot run the search.
- If known-relevant papers were supplied, say which terms should retrieve each
  one — that is the cheapest sanity check available.

Example of the standard I want
Weak:   "social media AND teenagers AND anxiety"
Strong: "(\"social media\" OR \"social networking\" OR Instagram OR TikTok)
         AND (adolescen* OR teen* OR youth) AND (anxiet* OR \"anxiety disorder\")"

Output
A string per database with syntax notes · the concept blocks · terms flagged
for verification · how to test recall against the known papers.

4. PRISMA-compatible protocol outline

Use when you need the protocol written before screening starts, not after.

Role: Methodologist drafting a protocol for registration. You write what will
actually be done, in the order it will be done.

Context
- Review question: [paste it]
- Criteria: [inclusion/exclusion]
- Search strategy: [databases and dates]
- Team: [how many screeners, how disagreements resolve]
- Planned synthesis: [narrative / meta-analysis / thematic]
- Registration target: [PROSPERO, OSF, or "internal only"]

Task
Draft the protocol outline against PRISMA-P.

Rules
- Cover every PRISMA-P section. Where I haven't given you the information,
  write [TO SPECIFY] with a note on what's needed — never fill a gap with a
  plausible-sounding default.
- Screening, extraction and appraisal each need: who does it, whether it's
  duplicated, and how conflicts resolve. Protocols usually fail here.
- Include the deviation policy — what happens when the search returns ten
  times what you expected.
- Do not cite the PRISMA statement with a specific year, author list or DOI
  unless I gave it to you.

Example of the standard I want
Weak:   "Two reviewers will screen studies."
Strong: "Two reviewers screen all titles/abstracts independently in Rayyan;
         disagreements go to a third reviewer; agreement reported as Cohen's
         kappa."

Output
Protocol outline by PRISMA-P section · [TO SPECIFY] gaps listed together ·
the three decisions to settle before screening opens.

5. Conceptual map from provided papers

Use when you have a stack of papers and want the shape of the field before writing.

Role: Research synthesist who maps what a body of work actually contains
rather than what a topic is usually assumed to contain.

Context
- Papers: [paste abstracts, or titles with key findings — label each P1, P2 …]
- Review question: [paste it]
- Field: [discipline]

Task
Build a conceptual map of the papers I gave you.

Rules
- Work only from the supplied papers. Every node and link traces to a paper ID.
  Do not add work you know of that I didn't provide — that silently changes
  the corpus the map claims to describe.
- Group by construct and by relationship, not by publication date.
- Show where papers use different names for the same construct, and the reverse:
  the same term used for different things. That distinction is usually the most
  valuable output.
- Mark thin areas as thin in this corpus, not as gaps in the field. You can only
  see what I sent.

Example of the standard I want
Weak:   "Several studies examine engagement."
Strong: "P2, P5, P9 all measure 'engagement' — P2 as time-on-platform, P5 as
         posts per week, P9 as a self-report scale. Not comparable."

Output
Concept clusters with paper IDs · relationships between them · terminology
conflicts · sparse areas in this corpus · what a map of the full field would
need that this one lacks.

This prompt maps the field structure from your actual evidence set. Do not run it without pasted content — asking for a conceptual map of a topic without provided papers will produce a hallucinated landscape.

6. Keyword gap check

Use when you want to know what your search string is likely missing before you rely on it.

Role: Search strategist who assumes every first-draft query has a blind spot
and looks for it deliberately.

Context
- Current search string: [paste it]
- Review question: [paste it]
- Papers retrieved: [titles, or a summary of what came back]
- Papers I expected but didn't get: [if any — these are the strongest clue]
- Field: [discipline, since vocabulary is field-specific]

Task
Find the terminology gaps.

Rules
- Check for: discipline-specific synonyms, older terms for the same concept,
  regional spelling, acronyms and their expansions, and adjacent fields that
  study the same thing under another name.
- If expected papers were missed, work backwards from them — which term would
  have caught each one.
- Propose additions as concrete string edits, and say for each whether it will
  widen recall usefully or mostly add noise.
- Do not estimate how many results a change returns. You cannot run it.

Example of the standard I want
Weak:   "Consider adding more synonyms."
Strong: "Pre-2010 work calls this 'computer-mediated communication'. Add
         (\"computer-mediated communication\" OR CMC) or you lose the
         foundational decade."

Output
Missing term categories with examples · the revised string · expected effect
of each addition · what to re-run and check.

How does AI help with summarizing research papers?

Summarization is the highest-volume task in a literature review: extracting consistent information from dozens or hundreds of papers. The core problem with unstructured summarization is inconsistency — different information captured per paper makes cross-study comparison impossible later. These six prompts enforce a consistent extraction structure you can reuse across your full paper set.

Critical rule for all six prompts: paste the actual text from the paper. Do not reference a paper by title or author and ask the model to summarize it. The model will draw on training data or hallucinate content it does not have. Your paste is the evidence.

7. Structured abstract extraction

Use when you need consistent fields out of a paper for an extraction table.

Role: Data extractor. You transcribe what a paper reports; you do not
interpret, complete or improve it.

Context
- Paper: [paste the abstract or full text]
- Fields I need: [population, design, sample size, measures, findings, limits]
- Review question: [so relevance can be judged]

Task
Extract a structured abstract in the fields listed.

Rules
- Every field comes from the text. If the paper does not report something,
  write "not reported" — never infer it from convention, and never estimate.
- Quote or closely paraphrase the paper's own wording for findings. Do not
  round, restate more strongly, or convert a null result into a trend.
- Keep the paper's hedging. "May be associated with" must not become "is
  associated with".
- Record effect sizes and confidence intervals exactly as printed, with units.
- Note where the abstract and the full text disagree, if you have both.

Example of the standard I want
Weak:   "The study found social media use increases anxiety."
Strong: "Reports a small positive association (r = .18, 95% CI [.09, .27]);
         authors state causality cannot be inferred from the design."

Output
The fields as a labelled list, with "not reported" where applicable, and a
note on anything ambiguous in the source.

8. Plain-language summary

Use when you need to explain a paper to someone outside the field without distorting it.

Role: Science communicator who refuses to buy clarity with accuracy.

Context
- Paper: [paste it]
- Audience: [policymaker, clinician, student, general public]
- What they'll do with it: [decision, coursework, background]
- Length: [how long]

Task
Write a plain-language summary.

Rules
- Keep the uncertainty. If the finding is correlational, the summary says so in
  plain words — "went together with", not "caused".
- Keep the population. A result in 200 undergraduates is not a result about
  people, and dropping that is the most common distortion in science writing.
- Define jargon on first use, or replace it. Don't do both.
- Include what the study could not show. A summary without limits reads as
  stronger evidence than it is.
- No numbers that aren't in the paper.

Example of the standard I want
Weak:   "Exercise reduces depression."
Strong: "Among 300 adults already in treatment, those assigned to 12 weeks of
         supervised exercise reported slightly fewer symptoms than those who
         weren't. It doesn't tell us whether exercise alone would help someone
         not in treatment."

Output
The summary · one line on what it does not show · any term you had to simplify
and what precision that cost.

9. Methods section audit

Use when you need to judge whether a study's design can support the claim it makes.

Role: Methodologist reviewing a methods section. You separate what the design
can support from what the authors assert.

Context
- Methods section: [paste it]
- The paper's main claim: [what they conclude]
- Field conventions: [discipline]
- My review question: [so relevance is clear]

Task
Audit the methods against the claim.

Rules
- State what this design can and cannot establish. Cross-sectional data cannot
  support a causal claim regardless of how the discussion is worded.
- Check: sampling and who it excludes, measure validity, whether analyses look
  pre-specified, handling of missing data, and control for obvious confounders.
- Separate "this is a flaw" from "this is a normal limitation of the design".
  Treating every constraint as a flaw makes the appraisal useless.
- Where the methods are too thin to judge, say what's missing instead of
  assuming best or worst practice.

Example of the standard I want
Weak:   "The sample size was small."
Strong: "n = 42 across three groups. With the reported effect size the study is
         underpowered for the interaction they interpret; that analysis should
         be read as exploratory."

Output
What the design supports · what it cannot · specific concerns with severity ·
what's missing from the reporting · how much weight to give this in synthesis.

10. Key claims extraction for a comparison table

Use when you're building a comparison table and need each paper's claims in comparable form.

Role: Extractor preparing rows for a cross-study table. Consistency across
papers matters more than elegance within one.

Context
- Paper: [paste it]
- Comparison dimensions: [the columns your table uses]
- Other papers already extracted: [so the format matches]
- Review question: [what the table is answering]

Task
Extract this paper's key claims against the dimensions.

Rules
- One claim per row, each with the evidence the paper offers for it and where
  it appears.
- Grade claim strength by what supports it: primary finding, secondary analysis,
  post-hoc, or authors' speculation in the discussion. Discussion speculation
  gets mistaken for finding constantly, so label it.
- If the paper says nothing on a dimension, write "not addressed" rather than
  leaving a blank a later reader will misread as a null result.
- Preserve units, timeframes and comparison groups. "40% higher" without a
  baseline is not extractable.

Example of the standard I want
Weak:   "Found a positive effect."
Strong: "Primary outcome: 12-week symptom reduction, d = 0.34 vs waitlist.
         Secondary (exploratory): no dose-response."

Output
Claims table: claim, evidence type, strength, location in paper · dimensions
not addressed · anything that won't compare cleanly with the other rows.

11. Contradictory findings flag

Use when your papers disagree and you need to know whether it's real disagreement.

Role: Synthesist who treats a contradiction as something to explain, not
something to average away.

Context
- Papers: [paste findings, labelled P1, P2 …]
- Review question: [paste it]
- What I've noticed: [the apparent conflict, if you've spotted one]

Task
Find and explain the contradictions across these papers.

Rules
- First check the conflict is real. Different populations, measures,
  timeframes or comparison groups produce apparent disagreement between
  studies that are both right about different things.
- For each genuine contradiction, list the candidate explanations: sample
  differences, measurement, analysis choices, publication bias, or a genuine
  moderator nobody has modelled.
- Do not resolve a contradiction by picking the larger or newer study. Say what
  evidence would resolve it.
- Only use the papers I gave you. Do not bring in outside findings.

Example of the standard I want
Weak:   "Studies show mixed results."
Strong: "P3 and P7 disagree, but P3 measured frequency and P7 measured
         duration. Not contradictory — they measure different exposures."

Output
Apparent vs genuine contradictions · candidate explanations for each · what
would settle it · how to report this honestly in the synthesis.

12. Verbatim quote extraction for thematic coding

Use when you're coding qualitative data and need quotes you can trust to be verbatim.

Role: Qualitative research assistant supporting thematic coding. Fidelity to
the source is the whole job.

Context
- Source: [paste the paper, transcript or excerpt]
- Coding frame: [themes you're working with, or "inductive — derive them"]
- Research question: [what the coding serves]

Task
Extract quotes relevant to the frame.

Rules
- Verbatim only. Character for character, including grammar, hesitations and
  emphasis. Never tidy a quote — a cleaned quote is a fabricated quote.
- Give the location of each: page, section, line, or participant ID.
- Where a quote needs surrounding context to be read fairly, include it in
  brackets rather than extending the quotation marks.
- Do not merge two passages into one quotation.
- If nothing in the source supports a theme, say so. An empty theme is a
  finding.
- Flag any quote whose meaning depends on context you'd have to supply.

Example of the standard I want
Weak:   The participant felt unsupported by management.
Strong: "I mean, they said they'd help but... nobody came." (P4, p.7)

Output
Quotes grouped by theme, each with location · themes with no support ·
context-dependent quotes flagged.

How do you build a research comparison matrix with AI?

A comparison matrix shows patterns across studies — which findings replicate, which contradict, which populations are missing. Building one manually across 20 papers takes a day. With consistent extraction from prompts 7–12 as input, structured comparison prompts compress that to an hour, with verification still on you.

13. Cross-study comparison matrix

Use when you have extracted several papers and need them side by side.

Role: Synthesist building a comparison matrix. You make differences visible
rather than smoothing them into a tidy grid.

Context
- Papers: [paste extractions, labelled P1, P2 …]
- Dimensions to compare: [or "derive them from the papers"]
- Review question: [what the matrix answers]

Task
Build the cross-study comparison matrix.

Rules
- Every cell traces to a paper. Empty means not reported — mark it, never
  leave it blank for a reader to misread as a null result.
- Do not convert between incomparable measures to fill a column. If P1 reports
  odds ratios and P4 reports mean differences, keep both and say they don't
  convert.
- Add a row for what each study measured, not just what it found. Two papers
  with the same outcome name often measured different things.
- Do not add papers I didn't give you.

Example of the standard I want
Weak:   | P1 | Positive effect | Large sample |
Strong: | P1 | d=0.34 (12wk, vs waitlist) | n=210, undergrads, US |

Output
The matrix · a note on which columns are genuinely comparable · rows where
the heterogeneity is too large to pool.

14. Thematic cluster analysis

Use when you have many papers and need the themes that actually run through them.

Role: Thematic analyst. You build themes from the corpus rather than sorting
the corpus into themes you brought with you.

Context
- Papers: [paste findings, labelled]
- Review question: [paste it]
- Existing framework: [if you're testing one, or "inductive"]

Task
Cluster these papers into themes.

Rules
- Each theme needs a definition and the paper IDs supporting it. A theme with
  one paper is an observation; label it as such.
- Report papers that fit no theme rather than forcing them. Those are usually
  where the interesting work is.
- Where a paper supports two themes, say so — overlap is information about the
  constructs, not a filing error.
- Distinguish a theme in the literature from a theme in your reading of it.
- Only the supplied papers.

Example of the standard I want
Weak:   "Theme: Barriers to adoption (P1, P3, P6, P9)"
Strong: "Theme: Institutional rather than individual barriers — cost and
         procurement, not skill or willingness (P1, P3, P6). P9 is adjacent
         but studies individual attitudes; listed separately."

Output
Themes with definitions and paper IDs · unassigned papers · overlaps ·
themes resting on a single study.

15. Agreement and disagreement table

Use when you need to state clearly where the literature converges and where it doesn't.

Role: Evidence synthesist. You report the spread of findings, not the average
of them.

Context
- Papers with findings: [labelled]
- Review question: [paste it]
- Study characteristics: [design, sample, setting per paper if available]

Task
Build an agreement and disagreement table.

Rules
- Three columns: consistent findings, contested findings, and single-study
  findings with no replication. That third column is usually the one people
  quote as if it were settled.
- For contested rows, note what differs between the studies — design, sample,
  measure, or era.
- Weight by study quality where you can, and say how you weighted. Do not
  count votes; five weak studies do not outweigh one strong one.
- Do not describe something as established from a single study.

Example of the standard I want
Weak:   "Most studies agree the intervention works."
Strong: "Consistent (4/5 studies, all RCTs): short-term symptom reduction.
         Contested: whether it holds at 6 months — P2 yes (n=400), P5 no
         (n=88, high attrition)."

Output
The table · what's genuinely settled · what's contested and why · what rests
on one study · confidence per row.

16. Population coverage analysis

Use when you want to know who the evidence base actually covers.

Role: Equity-minded reviewer. You check who the literature studied before
treating its conclusions as general.

Context
- Papers: [with sample descriptions, labelled]
- Review question: [and the population it implies]
- Target population: [who the conclusions are meant to serve]

Task
Analyse population coverage.

Rules
- Tabulate the reported sample characteristics per study: age, sex/gender, race
  and ethnicity, geography, socioeconomic context, clinical status. Mark
  "not reported" separately from "not represented" — they are different
  failures and both matter.
- Name the mismatch between the studied populations and the target population
  explicitly.
- Note over-representation too. A literature built on university students or
  on one country's health system generalises poorly, and it's rarely stated.
- Do not estimate demographics a paper didn't report.

Example of the standard I want
Weak:   "Samples were fairly diverse."
Strong: "5 of 6 studies were US undergraduate samples (median age 20). No
         study reported income. Conclusions about working adults are
         unsupported by this evidence base."

Output
Coverage table · populations absent · populations over-represented · what this
limits the review from claiming · how to word that limitation.

17. Operationalization comparison

Use when papers use the same word and you suspect they don't mean the same thing.

Role: Measurement-focused reviewer. You compare how constructs were
operationalised before comparing results.

Context
- Papers: [with their measures and definitions, labelled]
- Construct: [the concept in question]
- Review question: [paste it]

Task
Compare how each paper operationalised the construct.

Rules
- For each: the conceptual definition, the instrument or measure, the items or
  units, and the timeframe.
- Say explicitly which studies are measuring the same thing and which only
  share a label. This is the most common reason a synthesis quietly goes wrong.
- Note validation status where reported, and flag ad-hoc single-item measures.
- Where measures differ enough that pooling would be invalid, say so plainly.
- Do not assume an unnamed instrument.

Example of the standard I want
Weak:   "All studies measured wellbeing."
Strong: "P1 used WEMWBS (14-item, validated). P4 used one item: 'How happy
         are you today?' (1-10). These should not be pooled."

Output
Operationalisation table · which studies are comparable · which share only a
name · implications for synthesis.

18. Replication status check

Use when you need to know how much of the evidence base has been independently confirmed.

Role: Reviewer assessing replication. You distinguish repeated findings from
repeated citations of one finding.

Context
- Papers: [labelled, with authorship, samples and findings]
- Key findings to check: [the claims that matter]
- Field norms: [replication culture varies enormously by discipline]

Task
Assess replication status of the key findings.

Rules
- Classify each: direct replication, conceptual replication, extension, or
  single untested finding.
- Check independence. Overlapping author teams, or several papers on one
  dataset, are not independent replications — this is routinely missed and it
  inflates apparent support.
- Flag findings that are widely cited but rest on one study.
- Note failed or null replications with the same weight as successful ones.
- Do not assume replication exists because a finding is well known.

Example of the standard I want
Weak:   "This finding is well replicated."
Strong: "Four papers report it, but P2, P3 and P6 analyse the same 2019 cohort
         and share two authors. One independent replication (P8, null)."

Output
Findings with replication status and independence check · non-independent
clusters · widely-cited single-study findings · confidence per finding.

How do you identify research gaps using AI prompts?

Gap analysis is the section that distinguishes a strong literature review from a summary. It requires holding the whole evidence base in mind and reasoning about what is missing, contradicted, or untested. These six prompts apply consistent gap-detection frameworks to your evidence set.

19. Contradiction and inconsistency scan

Use when you want the inconsistencies in the evidence base surfaced deliberately.

Role: Critical reviewer whose job is to find where the literature does not
hang together.

Context
- Papers: [labelled, with findings and methods]
- Review question: [paste it]
- Field: [discipline]

Task
Scan for contradictions and inconsistencies.

Rules
- Look for: findings that conflict, methods that conflict with stated aims,
  conclusions that outrun the data, and definitions that shift between papers.
- Include within-paper inconsistency — an abstract stronger than the results,
  a discussion that quietly drops a null primary outcome.
- For each, state whether it's a real conflict or an artefact of different
  designs.
- Rank by how much it should change the review's conclusions.
- Only the supplied papers, and no invented findings.

Example of the standard I want
Weak:   "Some inconsistencies were noted."
Strong: "P4's abstract reports 'significant improvement'; the results give
         p = .07 on the primary outcome, significant only on a secondary
         measure. The abstract overstates it."

Output
Inconsistencies ranked by consequence, each with type, evidence and whether
it's genuine · which change the review's conclusions.

20. Understudied population write-up

Use when you've found a population gap and need it written up defensibly.

Role: Reviewer writing a gap statement that survives peer review — which
means a gap in the evidence you searched, not the evidence that exists.

Context
- Population coverage findings: [what you found]
- Search strategy: [what you searched and how — this bounds the claim]
- Review question: [paste it]
- Why this population matters: [the reason it's worth naming]

Task
Write the understudied-population section.

Rules
- Scope the claim to your search. "No studies in our search examined X" is
  defensible; "no studies have examined X" is not, and a reviewer will find
  the exception.
- Name what would fill the gap: design, sample, outcome.
- Explain why the gap matters for the review's question — an unexplained gap
  reads as padding.
- Consider why it exists: funding, access, historical exclusion, or the term
  not being searchable.
- Do not cite work you weren't given.

Example of the standard I want
Weak:   "More research is needed on older adults."
Strong: "No study in our search (5 databases, 2015-2025, English) included
         participants over 65, though the intervention is delivered mainly in
         settings serving that group."

Output
The section, scoped to the search · why it matters · what would fill it ·
likely reasons it exists.

21. Methodological gap analysis

Use when the gap is in how the field studies something, not who it studies.

Role: Methodologist identifying where a field's methods limit what it can
learn.

Context
- Papers with designs: [labelled]
- Review question: [paste it]
- Field conventions: [what's normal here]

Task
Identify the methodological gaps.

Rules
- Map designs used against designs the question needs. If every study is
  cross-sectional, the field cannot answer a causal question no matter how many
  papers accumulate — say that directly.
- Cover: design types, measurement (self-report vs observed), timeframes,
  comparison groups, preregistration and open data.
- Distinguish a gap that's a genuine opportunity from one that exists because
  the stronger design is impractical or unethical here.
- Say what a study filling the gap would look like, concretely.
- No invented citations.

Example of the standard I want
Weak:   "More longitudinal research is needed."
Strong: "All 7 studies are cross-sectional with self-report at one timepoint.
         The directionality question is unanswerable from this design class;
         a 12-month panel with two measures would settle it."

Output
Design coverage vs what the question needs · gaps ranked by consequence ·
which are opportunities vs constraints · specification of the study that fills
the most important one.

22. Specific future research questions

Use when you're writing the future-research section and don't want it to read as filler.

Role: Reviewer writing future directions someone could actually act on.

Context
- Gaps identified: [from your analysis]
- Field: [discipline]
- Feasibility constraints: [what's realistic — funding, access, ethics]
- Who might do this work: [audience for the section]

Task
Write specific future research questions.

Rules
- Every question names population, design, and outcome. "More research is
  needed" is the phrase this section exists to avoid.
- Tie each to a gap you actually identified. A question with no gap behind it
  is padding.
- Order by tractability, not by how interesting they sound, and note what each
  would take.
- Flag any that are ethically or practically hard, and say why.

Example of the standard I want
Weak:   "Future work should explore long-term effects."
Strong: "Does the 12-week effect persist at 12 months in adults over 65? A
         two-arm RCT with follow-up at 3, 6 and 12 months; primary outcome
         the same validated scale used by P1 and P3 so results pool."

Output
Questions ordered by tractability, each with design, population, outcome and
the gap it addresses · the two most important · what each requires.

23. Theoretical framework gaps

Use when you suspect the field is measuring things without a theory connecting them.

Role: Theory-oriented reviewer. You check whether a field's findings sit
inside an explanatory framework or just accumulate.

Context
- Papers: [with any theoretical framing, labelled]
- Review question: [paste it]
- Field's dominant theories: [if you know them]

Task
Analyse the theoretical framework gaps.

Rules
- For each paper, name the theory it invokes — and whether it's actually used
  to derive predictions or just cited in the introduction. That distinction is
  the finding here more often than not.
- Identify findings no current framework explains, and frameworks that predict
  things nobody has tested.
- Note atheoretical work explicitly; in some fields it's the majority.
- Do not invent a theory or attribute one to a paper that didn't name it.

Example of the standard I want
Weak:   "Studies lack theoretical grounding."
Strong: "5 of 8 cite Self-Determination Theory in the introduction; only P2
         derives a hypothesis from it. The rest use it decoratively."

Output
Theory use per paper, decorative vs load-bearing · unexplained findings ·
untested predictions · where the field needs theory rather than more data.

24. Implications gap for a specific audience

Use when the evidence exists but nobody has translated it for the people who'd use it.

Role: Knowledge translation specialist. You look for the gap between what is
known and what is usable.

Context
- Findings: [what the literature establishes]
- Target audience: [clinicians, teachers, policymakers, practitioners]
- What they decide: [the actual decisions this evidence should inform]
- Current guidance: [what they're working from now, if you know]

Task
Identify the implications gap for this audience.

Rules
- Separate: findings ready to act on, findings needing translation, and
  findings too preliminary to act on. Practitioners routinely act on the third
  category because nobody labelled it.
- For actionable findings, state the specific decision they inform.
- Name what's missing to make the rest usable — effect sizes in meaningful
  units, cost data, implementation detail, or evidence in their setting.
- Do not overstate readiness to be helpful.

Example of the standard I want
Weak:   "Practitioners should consider these findings."
Strong: "Ready to act on: the 12-week protocol (3 RCTs, consistent). Not ready:
         dosage — all studies used one schedule, so 'more is better' is
         untested and shouldn't be inferred."

Output
Three tiers of readiness · decisions each supports · what's missing for the
rest · how to phrase the uncertainty without paralysing the reader.

How do you use AI to critique research methodology?

Methodology critique applies consistent quality standards across papers that describe their methods with varying levels of detail. These six prompts apply established frameworks — risk of bias criteria, CASP checklists — to what you paste in, producing per-paper quality notes you can aggregate into an evidence quality summary.

25. Risk of bias checklist

Use when you're appraising quantitative studies and need it consistent across the set.

Role: Methodologist conducting risk-of-bias assessment. You judge reporting
as well as conduct, and you keep them apart.

Context
- Study: [paste methods and results]
- Design: [RCT, cohort, case-control, cross-sectional]
- Tool: [Cochrane RoB 2, ROBINS-I, NOS, or "pick and justify"]
- Review question: [what the appraisal serves]

Task
Apply the risk-of-bias checklist.

Rules
- Judge each domain from what the paper reports. "Not reported" is its own
  rating — unclear risk, not low and not high. Conflating them is the standard
  error in appraisal.
- Quote or cite the text supporting each judgement. An unsupported rating
  isn't reproducible.
- Distinguish poor reporting from poor conduct, and say which you're rating.
- Give the overall judgement with the reasoning, and say which domain drove it.
- Do not assume standard practice was followed because the journal is reputable.

Example of the standard I want
Weak:   "Randomisation: unclear."
Strong: "Randomisation: unclear risk. States 'participants were randomly
         assigned' (p.4) with no sequence generation or concealment method.
         Reporting gap, not necessarily a conduct flaw."

Output
Domain-by-domain ratings with supporting text · overall risk and what drove
it · reporting vs conduct · what to request from authors.

26. CASP qualitative appraisal

Use when you're appraising qualitative work, where the quantitative checklists don't fit.

Role: Qualitative methodologist. You appraise against qualitative standards
rather than penalising a study for not being a trial.

Context
- Study: [paste it]
- Methodology: [grounded theory, phenomenology, ethnography, thematic]
- Review question: [what the appraisal serves]

Task
Apply a CASP-style qualitative appraisal.

Rules
- Judge methodological coherence: does the method fit the question, and does
  the analysis fit the method. That matters more than sample size here.
- Do not apply quantitative criteria. Small samples, non-random sampling and
  researcher involvement are features of the design, not flaws.
- Check reflexivity — is the researcher's position stated and considered.
- Assess whether interpretation is grounded in data: are there enough quotes,
  and do they support the claims made from them.
- Where reporting is too thin to judge, say so rather than assuming.

Example of the standard I want
Weak:   "Sample size was only 12 participants."
Strong: "12 participants is appropriate for IPA. Saturation isn't discussed,
         which matters more here than n — worth noting as a reporting gap."

Output
CASP domains with judgements and supporting text · methodological coherence ·
reflexivity · overall confidence and what it rests on.

27. Sample size and power note

Use when you need to judge whether a study had a fair chance of detecting its effect.

Role: Statistical reviewer. You are careful about the difference between no
evidence of effect and evidence of no effect.

Context
- Study: [paste methods, sample size, results]
- Reported power analysis: [paste it, or "none reported"]
- Effect sizes: [observed, with CIs]
- Design: [and number of groups or comparisons]

Task
Write a sample size and power note.

Rules
- If a power analysis is reported, check it was for the primary outcome and
  that the assumed effect size was justified rather than reverse-engineered
  from a feasible n.
- Never compute post-hoc power from the observed effect. It's a known fallacy
  and adds nothing beyond the confidence interval — say so if I ask for it.
- Read the confidence interval instead. A wide CI spanning null and a large
  effect means uninformative, not "no effect".
- Check power for the analysis actually interpreted — interactions and subgroups
  need far more than main effects, and this is where most claims break.
- Note multiple comparisons without correction.

Example of the standard I want
Weak:   "The study was underpowered."
Strong: "Powered for the main effect (n=120, d=0.5), but the paper's headline
         is a subgroup interaction. CI [-0.1, 0.9] — consistent with anything
         from null to large."

Output
Power assessment for the interpreted analysis · what the CIs support · whether
null results are informative · how much weight this study should carry.

28. External validity note

Use when you need to say where a finding does and doesn't transfer.

Role: Reviewer assessing generalisability. You keep sampling and setting in
view when a finding gets stated as a general truth.

Context
- Study: [paste sample, setting, procedure]
- The claim: [what's being generalised]
- Target context: [where you want to apply it]
- Other studies: [if you're comparing across settings]

Task
Write an external validity note.

Rules
- Compare studied population to target on the characteristics that plausibly
  moderate the effect — not every difference matters, and listing all of them
  hides the ones that do.
- Consider setting and procedure too. Lab conditions, intensive support, or
  unusually motivated volunteers all limit transfer independently of who was
  studied.
- Note selection into the study: volunteers, clinic attenders and completers
  differ systematically from the population.
- State what would need to be true for the finding to transfer.
- Don't dismiss a study as ungeneralisable without naming the specific threat.

Example of the standard I want
Weak:   "Results may not generalise."
Strong: "Delivered by the developers in a research clinic with weekly
         supervision. Transfer to routine practice depends on whether the
         effect survives non-specialist delivery — untested here."

Output
Studied vs target comparison on moderating characteristics · specific threats ·
conditions for transfer · how confidently the claim can be applied.

29. Cross-study quality summary table

Use when you've appraised the studies and need the quality picture in one place.

Role: Synthesist summarising quality across an evidence base so a reader can
weigh the conclusions.

Context
- Appraisals: [your per-study results]
- Studies: [labelled, with designs]
- Review question: [paste it]
- Tool used: [which appraisal instrument]

Task
Build the cross-study quality summary table.

Rules
- Rows are studies, columns are domains, cells carry the rating with a short
  reason. A table of colours with no reasons can't be checked.
- Add a row summarising each domain across studies — a weakness shared by every
  study is a property of the field and should be reported as one.
- Say how quality affects the review's conclusions: which findings rest on
  low-risk studies and which rest on high-risk ones.
- Do not average ratings into a single score. It hides exactly what the table
  exists to show.

Example of the standard I want
Weak:   "Overall quality was moderate."
Strong: "Blinding: high risk in 6 of 7 — a field-level limitation, not a
         study-level one. The main finding rests on the one low-risk study."

Output
Quality matrix with reasons · domain-level patterns · which conclusions rest
on which quality tier · how to state this in the results.

30. AI use disclosure for your methods section

Use when you used AI in the review and need to disclose it accurately.

Role: Research integrity advisor. You write disclosures that are specific
enough to be reproducible and honest enough to survive scrutiny.

Context
- Where AI was used: [screening, extraction, summarising, drafting, coding]
- Tool and version: [model name and version, and the date range of use]
- Human oversight: [what was checked, by whom, and how much]
- Target: [journal, thesis, institution — and their policy if you have it]

Task
Draft the AI use disclosure.

Rules
- Be specific about each stage. "AI was used to assist" tells a reader nothing
  and is the version most likely to be challenged.
- State the verification for each use — the disclosure's weight comes from what
  a human checked, not from admitting AI was involved.
- Do not claim oversight that didn't happen. If nothing was independently
  verified at some stage, say that.
- Name the tool and version. Behaviour differs between versions, which is a
  reproducibility issue.
- If I haven't given you the venue's policy, write to the strictest common
  standard and flag that it needs checking against [VENUE POLICY].

Example of the standard I want
Weak:   "AI tools were used to assist with this review."
Strong: "[Model, version] was used to extract predefined fields from included
         full texts (Jan-Mar 2026). All extractions were checked against source
         papers by two authors; 4 of 210 fields required correction. No AI was
         used for screening decisions or for drafting the results."

Output
The disclosure statement · a per-stage table of use and verification · the
methods-section sentence · what to confirm against the venue's policy.

Prompt 30 is the one most researchers forget. Journals increasingly require disclosure of AI assistance, and having a prompt that generates a consistent disclosure statement means you document your process while the work is fresh — not reconstruct it at submission. For a full guide to versioning and documenting your prompt workflow as part of your research method, see building a reproducible AI research workflow.

How do you get more from these 30 prompts?

Three practices separate output you can use in a submitted paper from output you need to rewrite.

  1. Paste actual text, every time. The most common failure mode is referencing a paper by title rather than pasting its content. The model draws on training data, mixes it with hallucinated details, and produces a summary that sounds plausible but is wrong. Full abstract plus methods section is the minimum paste for reliable extraction.
  2. Work in batches of five to ten papers. Context windows have grown, but per-paper attention degrades across very large pastes. Summarize in batches, verify each batch, then synthesize across batch summaries. This also creates natural checkpoints for citation verification.
  3. Save your extraction protocol. Once you have a structured extraction prompt that produces output in exactly the format your review matrix needs, save it. A consistent protocol run across every paper in your review is what makes AI-assisted synthesis defensible in a methods section — and what makes the next review faster.

For downstream work — efficiently summarizing a large paper set and maintaining citation integrity throughout — see how to summarize 50 papers without losing citations. For applying this prompt set to grant writing, prompt templates for grant writing and abstracts covers the funder-specific translation.

How Prompt Architects fits this workflow

All 30 prompts above work in any AI tool you already use. What Prompt Architects adds is the infrastructure that makes them repeatable across projects: a prompt library where you save your extraction protocol, comparison matrix template, and gap analysis prompts with your research question and field context already stored.

The Contexts feature lets you store your research question, inclusion criteria, and field scope once, so that context injects automatically into each prompt run rather than requiring manual copy-paste at the start of every session. When you begin a new chapter or a new systematic review, you open your saved library and run your protocol — not rebuild it.

Prompt Architects is free to start, no credit card required. The /ai-for-researchers landing page shows how the library and Contexts features map to the full literature review workflow.


Pick the five prompts that match your current bottleneck, run them on your next batch of papers, and save any extraction template that produces directly usable output. The protocol you standardize today is the one that makes your methods section defensible at submission.

Start free — save your first research extraction prompt in under two minutes →

Frequently asked questions

Free Chrome Extension

Stop rewriting prompts. Start shipping.

Works with ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3 & Kling. 5.0★ on the Chrome Web Store.

Create An Account
NH
Written by
Nafiul Hasan
Founder, Prompt Architects

Nafiul Hasan is the founder of Prompt Architects — an all-in-one AI prompt generator, enhancer, and library Chrome extension for ChatGPT, Claude, Gemini, Grok, Midjourney, Ideogram, Veo3, and Kling. He writes about prompt engineering, AI workflows, structured outputs, and what actually works when building production AI tools. Prior to founding Prompt Architects he worked across web product and developer tooling.

Prompt engineeringAI workflows (RAG, agents, structured output)Chrome extensionsIndie SaaSMulti-LLM tooling