From a Broad Topic to PICO/PECO: Using AI to Frame a Research Question Without Fabricated Sources

Briefly and clearly: artificial intelligence can help decompose a topic, compare alternative formulations and expose missing design elements. It cannot independently establish novelty, prove that no prior studies exist, or guarantee bibliographic accuracy. A research gap must be supported by a transparent, reproducible search, critical appraisal of primary sources and a realistic feasibility assessment.

[Approximate reading time: 12 minutes]

Colleague, welcome.

A broad topic is not yet a research question. A phrase such as “to investigate biomarkers in cardiovascular disease” does not specify who will be studied, which intervention or exposure will be evaluated, what the comparator is, which outcome matters most, or when it will be measured. It therefore cannot adequately anchor a protocol, sample-size calculation or statistical analysis plan.

A well-formulated question, by contrast, guides eligibility criteria, literature searching, data collection, analysis and interpretation. This is where AI can be useful—as a tool for disciplined intellectual structuring. It becomes methodologically unsafe when asked to “find the gap”, “prove relevance” or provide ready-to-cite sources without independent verification.

In a 2023 empirical study of then-current GPT-3.5 and GPT-4 outputs, 55% and 18% of citations, respectively, were fabricated, and substantive errors also occurred in citations to real works. These percentages must not be extrapolated mechanically to current systems, web-enabled workflows or every discipline. The operational conclusion remains valid: verify every source, DOI, number and factual claim against the primary source.

1. Select the framework after identifying the question type

PICO for intervention questions

  • P — Population / Problem: the population or clinical problem;
  • I — Intervention: a treatment, diagnostic, preventive or rehabilitation technology;
  • C — Comparison: usual care, placebo, an alternative or another defensible comparator;
  • O — Outcome: a patient-important benefit, harm or other endpoint.

PECO for exposure questions

  • P — Population: a defined population;
  • E — Exposure: a biomarker, behavioural, occupational, environmental or other exposure;
  • C — Comparator: no exposure, a lower level or another exposure category;
  • O — Outcome: the outcome expected to be associated with the exposure.

Adding T — Time and S — Setting is often useful. PICO is not universal, however. Diagnostic-accuracy, prognosis, qualitative, implementation and complex-systems questions may require tailored frameworks. The correct order is therefore: identify the question type first; select the mnemonic second.

2. Convert the topic into a controlled specification

Before prompting an AI system, document what it does not know:

  1. Which clinical, managerial or scientific decision should the answer support?
  2. Which patients, data, biospecimens, equipment, skills and time are actually available?
  3. Which primary outcome matters to patients, clinicians, commissioners or the health system?
  4. Which ethical, legal, confidentiality and contractual constraints apply?
  5. Is the target a causal effect, prediction, diagnostic accuracy, prevalence, association or lived experience?

Then complete a working matrix.

ComponentWhat must be specifiedPrimary data requiredWhat the literature must verify
Pinclusion and exclusion criteria; age; diagnosis; stage; contextrecords, registries and diagnostic confirmationwhether the boundaries are defensible and clinically useful
I / Eexact intervention protocol or exposure operationalisationdose, duration, device, laboratory method and timingstandardisation, reproducibility and biological plausibility
Cclinically and ethically valid comparatorcontrol-group data and allocation ruleswhether the contrast answers the intended question
Oone primary and a limited set of secondary outcomesvalidated scales, laboratory or instrumental measures and eventsvalidity, clinical importance and selective-reporting risk
T / Sfollow-up horizon and settingdates, duration, level of care and pathwaywhether the context matches the intended application

3. Give AI a bounded role

ROLE. Act as a methodological assistant in biomedical research. Generate structural alternatives, but do not certify facts or determine scientific novelty.

INPUT. Broad topic: [topic]. Context: [institution/population/country]. Available data and resources: [list]. Constraints: [time, equipment, budget, ethics, confidentiality]. Intended decision: [clinical/predictive/diagnostic/organisational].

TASK. Propose three materially different research questions. For each option: 1. justify PICO, PECO or another framework; 2. provide P/I/E/C/O/T/S in a table; 3. state a working hypothesis or target estimand; 4. list the variables and primary data required; 5. identify major confounders, biases and feasibility risks; 6. propose search concept blocks, but not a purportedly final strategy; 7. list the decisions that require human resolution before protocol approval.

RED LINES. Do not claim that the question is novel or that a gap has been proven. Do not invent references, DOIs, guidelines or statistics. Do not replace missing data with assumptions. Label every unverified statement REQUIRES VERIFICATION. Do not request or reproduce identifiable patient or confidential information.

OUTPUT. A comparative table, FINER ranking and a separate human-verification checklist.

This prompt does not generate evidence. It constrains idea generation and makes uncertainty visible.

4. Test the options against FINER and the real clinical base

Assess each option as:

  • Feasible: sufficient sample, events, time, budget, data access and competence;
  • Interesting: meaningful to the team and intended users;
  • Novel: adds knowledge rather than duplicating completed work;
  • Ethical: has an acceptable benefit–risk balance and meets current requirements;
  • Relevant: can influence decisions, practice, policy or subsequent research.

For observational studies, also define likely confounders, temporality, selection and information bias, and variables that should not be adjusted for indiscriminately. For intervention studies, examine the comparator, allocation concealment, blinding, adherence, harms and clinical—not merely statistical—importance.

5. Establish the gap through reproducible searching

Novelty is not a property of elegant wording. It is an evidence-based judgement made after mapping existing knowledge against the intended contribution.

A defensible workflow is:

  1. Check recent systematic or scoping syntheses, clinical guidelines and research-priority documents.
  2. Identify completed, ongoing and registered studies in relevant databases and registries.
  3. In PubMed, combine MeSH with free-text synonyms and inspect Search Details to understand how the query was translated.
  4. Depending on the question and access, complement PubMed/MEDLINE with Cochrane Library, Embase, Scopus, Web of Science and specialist resources. No single database is universally sufficient.
  5. Preserve the database and platform, exact strategy, date, limits and number of retrieved records. For systematic reviews, use PRISMA-S.
  6. Verify DOIs through the publisher or DOI registry and confirm bibliographic metadata using independent identifiers.
  7. Distinguish absence of evidence from evidence of absence. “I did not find it” does not mean “it has never been studied”.

A technical correction that prevents missed evidence

PICO structures the question and eligibility criteria, but it should not automatically become one search string containing every element joined by AND. Comparators and outcomes are often absent from titles and abstracts or poorly indexed. Requiring them may reduce sensitivity and exclude relevant studies. Search strategies should therefore use the principal concepts, balance sensitivity and precision, be tested against known key records and, where possible, be peer reviewed by an information specialist.

6. Lock the protocol, then select the reporting guideline

EQUATOR Network helps identify the appropriate reporting guideline: CONSORT for randomised trials, STROBE for observational studies, PRISMA for systematic reviews, STARD for diagnostic-accuracy studies, TRIPOD and others for their respective designs. A reporting checklist cannot repair a weak design after data collection. Use it during planning, but do not confuse it with a protocol, ethics review or proof of novelty.

AI use in manuscript production should be disclosed according to journal policy. ICMJE states that AI systems cannot be authors; human authors remain responsible for accuracy, integrity, originality, absence of plagiarism and complete attribution.

7. Do not put non-public material into a public chatbot

Without an institutionally approved secure environment, do not enter:

  • identified or potentially re-identifiable patient data;
  • full clinical records, identifiable images or genomic data;
  • unpublished manuscripts, grant applications, peer reviews or data covered by confidentiality obligations;
  • trade secrets, patent-sensitive solutions and partner information;
  • credentials, access keys or internal organisational documents.

The open-science principle “as open as possible, as closed as necessary” does not require disclosure of personal data or abandonment of legitimate commercialisation rights. Openness must be lawful, governed and planned.

8. A compact PECO example

Broad topic: biomarkers and incident heart failure in cardiometabolic comorbidity.

Structured option:

  • P: adults with hypertension and obesity but no heart failure at baseline;
  • E: an elevated baseline level of a pre-specified biomarker;
  • C: a lower or reference level of the same biomarker;
  • O: newly diagnosed heart failure during a defined follow-up period;
  • T: for example, 24 months;
  • S: a defined clinical network or registry with a documented follow-up pathway.

Before approval, a human team must assess event frequency, assay standardisation, serial-measurement availability, outcome adjudication, loss to follow-up, key confounders, existing risk models, incremental clinical value and external validation. AI can produce the checklist; it cannot demonstrate that the data exist or that the study is adequately powered.

9. Responsibility matrix

Appropriate AI supportHuman work requiredNever delegate
alternative formulations; PICO/PECO tables; gap spotting in the logic; synonym generation; formatting a search logfact and source verification; alignment with available data; bias assessment; protocol approval; ethics and statistical reviewauthorship accountability; novelty judgement; informed consent; confidentiality protection; clinical decisions; final interpretation

10. Pre-approval checklist

  • ☐ Question type and framework are explicit.
  • ☐ P/I/E/C/O/T/S are operationalised.
  • ☐ The primary outcome is important and measurable.
  • ☐ Available data and their quality are known.
  • ☐ Major biases and confounders are anticipated.
  • ☐ Existing and ongoing syntheses and studies have been checked.
  • ☐ Searching is reproducibly documented.
  • ☐ Every reference and DOI has been verified.
  • ☐ Ethics, confidentiality and data rights are addressed.
  • ☐ AI use is documented and responsibility remains human.

Conclusion

The proper role of AI at the start of a study is to accelerate thinking without replacing evidence. PICO or PECO converts a broad topic into a controlled specification; reproducible searching establishes what is already known; and methodologists, statisticians, clinicians and stakeholders decide whether the question is feasible, ethical, relevant and genuinely novel.

A clear process cannot guarantee protocol approval, publication or a successful defence. It materially reduces the risks of methodological error, fabricated sources, unusable data and costly late redesign.

Would you like to apply this workflow to your own topic? Register for the practical programme on effective and responsible use of AI in research: https://forms.gle/pdaNEssYrqP9oJxv8

For an individual methodological review, fundraising strategy or research-management plan, book a consultation: https://calendly.com/koleksii/1-hour

Contact me if I can help you or collaborate in any way.

Where do you draw the boundary between useful AI-assisted decomposition and the unacceptable delegation of scientific judgement?

Please share your view—or questions—in the comments.

I wish you inspiration, consistency and enough health for all of it.

Sincerely yours,
Oleksii Kalmykov
https://kalmykov.info/en

References

  1. Thomas J, Kneale D, McKenzie JE, Brennan SE, Bhaumik S. Chapter 2: Determining the scope of the review and the questions it will address. In: Cochrane Handbook for Systematic Reviews of Interventions, version 6.5. Cochrane; 2024.
  2. Lefebvre C, Glanville J, Briscoe S, et al. Chapter 4: Searching for and selecting studies. In: Cochrane Handbook for Systematic Reviews of Interventions. Current online version.
  3. Morgan RL, Whaley P, Thayer KA, Schünemann HJ. Identifying the PECO: A framework for formulating good questions to explore the association of environmental and other exposures with health outcomes. Environment International. 2018;121(Pt 1):1027–1031.
  4. EQUATOR Network. What is a reporting guideline?
  5. Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA Statement for Reporting Literature Searches in Systematic Reviews. Systematic Reviews. 2021;10:39.
  6. National Library of Medicine. PubMed Help: Advanced Search, Search Details and MeSH.
  7. Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports. 2023;13:14045.
  8. International Committee of Medical Journal Editors. Use of Artificial Intelligence by Authors. ICMJE Recommendations.
  9. World Medical Association. Declaration of Helsinki: Ethical Principles for Medical Research Involving Human Participants. 2024 revision.
  10. European Commission, European Research Executive Agency. Open science.

Leave a Reply

Your email address will not be published. Required fields are marked *

    Feedback form