Researchers can use AI safely when they assign it a bounded assistive role and keep human responsibility for evidence, methods, decisions, and claims. Useful roles include expanding search terms, ranking records for review, proposing structured extraction fields, drafting or explaining code, checking language, and organizing notes. Generated output is never a scholarly source. Read the original work, verify every citation and extracted fact, and retain the scripts, decisions, and evidence trail that support the final result.
The correct level of caution depends on the input and the consequence of an error. Public literature and low-risk language editing can often be used with documented review. Participant data, restricted datasets, unpublished manuscripts, grant applications under confidential review, and sensitive institutional material need prior approval and an environment that is expressly authorized for those data. For example, NIH prohibits generative AI in its peer-review process because submitting grant materials or critiques to online AI tools violates its confidentiality requirements. NIH peer-review notice
Use a red, amber, green decision rule before each task. Green work is public and reversible with a human check. Amber work is useful only with validation, a record of the tool and settings, and a researcher who can inspect the originals. Red work is prohibited or paused until the data steward, institutional review board, sponsor, funder, and venue rules say it is allowed. At publication, disclose material AI use as required, do not name AI as an author, and accept human accountability for the entire submission. ICMJE guidance
The operating principle
Use AI to create candidates for human research work, not to replace the evidence or the accountable decision. A candidate search term is not a search strategy. A candidate citation is not a citation. A candidate data field is not an extracted fact. A generated statistical explanation is not an analysis. The researcher must be able to point from every material statement, inclusion decision, number, code result, and conclusion back to an original source, dataset, method, or documented judgment.
This is not a rule against automation. It is a rule about evidence and accountability. A reference manager, script, spreadsheet formula, classifier, or language model can reduce clerical work. The research record must still make it possible for another qualified person to understand what was done, reproduce the relevant computation or decision, identify the inputs, and challenge an error.
Publisher guidance is consistent on the accountability boundary. ICMJE says authors must disclose use of AI-assisted technologies, review and edit generated content, ensure attribution, and must not cite AI-generated material as a primary source. It also says an AI tool cannot be an author because it cannot take responsibility for accuracy, integrity, or originality. ICMJE guidance COPE likewise states that authors remain fully responsible and should transparently describe which AI tool was used and how. COPE position
Red amber green decision matrix
Classify both the task and the data before entering anything into a tool. The same action can move from green to red when its inputs change from public abstracts to identifiable participant notes.
| Level | Suitable work | Required controls | Do not do this |
|---|---|---|---|
| Green | Query expansion from public concepts, outline alternatives, formatting, grammar, code explanation on non-sensitive examples, or drafting a screening checklist | Human checks the result before use; no confidential input; retain the final human decision | Treat suggested references, facts, or analyses as verified |
| Amber | Title and abstract ranking, structured extraction from accessible papers, coding assistance, analysis-plan critique, synthesis of approved documents, or language editing of a manuscript | Approved environment and data classification; validation sample; original-source checks; a log of tool, model, date, prompt or settings, and human review | Let the tool make an unreviewed inclusion, exclusion, fact, statistical, or authorship decision |
| Red | Identifiable participant data, controlled-access data, confidential peer-review material, unpublished grant or manuscript content where sharing is not authorized, restricted sponsor data, or any work the protocol or contract forbids | Stop and obtain written authorization from the appropriate data steward, institutional process, sponsor, funder, or editor before proceeding | Upload, paste, or summarize the material into a public or unapproved tool |
The matrix is deliberately conservative. “Amber” does not mean safe by default. It means there may be a defensible workflow after the data, provider, contract, approval, and validation controls are known. “Red” may later become amber only after an authorized secure environment and documented institutional approval exist. A generic promise that a tool is private is not by itself approval to use protected research material.
Practical uses with controls that preserve rigor
Discovery and query expansion
AI can help turn a research question into synonyms, related concepts, controlled-vocabulary candidates, alternative spellings, acronyms, and exclusion terms. Treat these as a brainstorming list. A librarian or researcher should turn them into the actual database query, inspect recall and precision, run the query in the canonical database, and save the exact string, date, database, filters, and result count.
Do not ask a general model to supply the final literature list and then cite what it names. It can omit key studies, merge papers, invent metadata, or offer a real title with the wrong author, year, or conclusion. Retrieve the original record and full text through a library, database, or publisher before adding it to a reference manager or evidence table. For systematic reviews, preserve the search and record flow expected by the relevant reporting standard. PRISMA-S provides a reporting checklist for literature searches, and the PRISMA flow diagram records the counts and reasons for inclusion and exclusion. PRISMA-S and PRISMA flow diagram
Title and abstract screening
An AI system can prioritize records that appear likely to meet prewritten criteria, flag possible duplicates, or identify records needing a second look. It should not silently decide eligibility. Freeze the protocol first, including population, intervention or exposure, comparison, outcomes, study design, language, date, and exclusion rules where relevant. Screen a representative pilot set manually, compare the tool's suggestions to independent human decisions, and define how disagreements are adjudicated.
Keep a screening log with the record identifier, reviewer, tool suggestion, final human decision, exclusion reason, and date. Audit a sample of both tool-prioritized inclusions and de-prioritized exclusions. This is especially important when missing an eligible study could bias a review. If the system cannot cite the exact title, abstract text, or full-text passage that motivated its suggestion, it has not produced a reviewable rationale.
Structured extraction
AI can turn a paper into candidate fields such as study design, population, intervention, outcome definition, sample size, effect estimate, limitation, or funding source. Require an evidence locator for every field, such as a page, table, figure, quotation, or section heading. A researcher then checks the locator against the original and records the verified value, not the model's paraphrase.
For high-impact extraction, use dual human extraction or a predefined validation sample and measure disagreement. Preserve the extraction schema, source file version, model output, reviewer correction, and final value. This makes it possible to assess whether the tool saves work without changing the evidence base. Never let a fluent summary conceal a missing confidence interval, a different outcome definition, a subgroup limitation, or a statement that belongs only in a discussion section.
Coding support
AI can explain a library error, draft a data-cleaning function, suggest test cases, translate a documented algorithm into another language, or help write comments and documentation. Keep the code in version control, review each change, run tests, and record the dependency versions, random seeds, operating environment, and input-data version needed to reproduce the result.
Do not paste credentials, private repository content, identifiable records, proprietary methods, or restricted data into an unapproved assistant. Do not accept generated code because it looks plausible. Run it on known test cases, inspect transformations, and compare results to a simple baseline. For statistical or scientific computation, preserve the executable script or notebook that produced each figure and table, not only a chat transcript that describes it.
Analysis support
AI is often useful as a critic of an analysis plan, a generator of questions for a methods meeting, an explainer of an output, or a draft source of code for a method the researcher already understands. It can also help construct a reproducible checklist for assumptions, missing data, multiplicity, model diagnostics, robustness checks, and visualization. The statistician or domain researcher must select the method, inspect the data, interpret the result, and decide whether the assumptions are met.
Do not ask a model to infer a result from a vague description and then write that result into a paper. Do not use generated synthetic values, images, or examples in a way that could be mistaken for observations. If synthetic data or images are used for a legitimate method, mark them clearly and document why and how they were created. NIH's current intramural research guidance similarly says synthetic data and images should be clearly marked so they are not treated as real. NIH research-conduct guidance
Language editing and communication
Language editing, plain-language summaries, alternative titles, formatting, and readability checks are often lower-risk uses when the text is authorized for the tool and the author verifies that meaning did not change. Check every changed number, qualifier, citation, and technical term. A wording assistant can accidentally turn “associated with” into “caused,” omit a limitation, or change the scope of an inference.
For manuscripts, follow the target journal or conference's current instructions before submission. Policies vary by venue and can change. If the tool materially drafted, rewrote, analyzed, or generated text, code, images, or data, disclose it at the required level of detail. Keep the disclosure factual: tool or provider, model or version when available, date, purpose, material inputs, settings where relevant, human validation, and whether any output remains in the final work.
A concrete literature-review workflow
Example
Hypothetical setup: a research team is preparing a review of interventions for a defined animal-health outcome. The protocol lists the review question, databases, date range, search concepts, inclusion and exclusion criteria, screening method, extraction fields, and planned synthesis. The team has approval to process public bibliographic records and lawful full texts in a specified environment, but not to upload unpublished peer-review material or participant-level data to external tools.
First, the team asks an approved tool for synonyms, species terms, intervention names, and possible controlled vocabulary. A researcher reviews this list with a subject librarian, creates the final search strings, runs them in named databases, and saves each query and export. The team deduplicates records using its reference-management process and records counts. The tool may rank titles and abstracts against the frozen criteria, but two reviewers make final inclusion and exclusion decisions on a pilot set and resolve disagreements using the protocol.
For included full texts, the tool creates a candidate extraction table only when every field includes a location in the original paper. Reviewers verify the location, correct the field, and preserve the source PDF version and the correction log. The narrative synthesis is written from the verified evidence table and cites the original studies, not the tool. Before submission, the team checks the venue policy, writes a disclosure, re-runs the search if required by the protocol, preserves the screening and extraction record, and ensures the PRISMA reporting material reflects the actual process.
The practical takeaway is that AI may reduce clerical effort while the research record remains independently auditable. If the tool cannot trace a suggestion to the original text, or if human review finds systematic missed studies or extraction errors, the team should narrow the tool's role rather than quietly accepting a less reliable review.
Confidentiality and approved environments
Data classification comes before prompt design. Ask the data steward or institutional security team whether the planned environment is authorized for the exact data class, whether the provider is an approved processor, what retention and training terms apply, where data may be processed, who can access logs, and whether the consent form, data-use agreement, protocol, funder terms, or sponsor contract allows the use. A local model is not automatically permissible, and a paid enterprise account is not automatically appropriate.
Do not assume that stripping obvious identifiers makes research data safe to upload. Re-identification risk, free text, combinations of fields, consent limits, contractual controls, and the purpose of the original collection may still matter. If human-participant data, protected health information, controlled access data, unpublished material, or proprietary sponsor information is involved, pause until the institutional review board, privacy office, data steward, or contract owner confirms the allowed workflow. NIH identifies human-subject protections, privacy, and unauthorized data disclosure as policy considerations for AI in research. NIH AI policy resource
There are explicit narrow prohibitions that illustrate the boundary. NIH prohibits reviewers from using generative AI to analyze or formulate critiques for NIH grant applications and R&D contract proposals. NIH peer-review notice NIH also states that sharing controlled-access human genomic data or derivatives with public generative tools violates the relevant non-transferability rules without appropriate written approval. NIH genomic-data notice These rules are not a universal legal test, but they show why a researcher's personal judgment about convenience cannot replace the applicable data and review requirements.
Preserve a reproducible AI use record
Record enough detail for a coauthor, reviewer, or future you to understand what role the tool played. The record can be a project log, protocol amendment, repository file, or methods appendix, depending on the work and policy. It should include:
- The research task, permitted data classification, approval basis, and the human responsible for the final decision.
- Tool or provider, model and version if exposed, account or environment type, access date, and material settings such as retrieval corpus, temperature, system instructions, or automation threshold.
- Prompt templates, input selection rules, tool outputs that influenced a decision, and the source documents or identifiers used to verify them.
- Screening decisions, extraction corrections, code changes, validation results, reviewer roles, and the final human-approved artifact.
- Known limitations, failure incidents, excluded uses, disclosure wording, and the applicable journal, funder, sponsor, or institutional policy version.
Do not turn a reproducibility log into a second disclosure channel for sensitive data. Store hashes, references, redacted prompts, or access-controlled records when raw content cannot safely be retained. The point is traceability, not indiscriminate preservation.
Bias, authorship, and publication integrity
AI can reproduce biases in its inputs, make weak generalizations look confident, and amplify a dominant framing by retrieving or summarizing more of the same kind of source. Counter this with a protocol that specifies inclusion criteria, search locations, population coverage, outcome definitions, and a plan to examine missing or underrepresented evidence. For qualitative research, retain the human codebook, reflexive notes, coding decisions, and adjudication record. A tool can propose labels, but it cannot substitute for the interpretive accountability of the research team.
AI is not an author, reviewer, guarantor, or source of evidence. Human authors must check accuracy, originality, attribution, conflicts, and disclosure. ICMJE says failure to disclose AI use may require corrective action and can be construed as misconduct in some circumstances. ICMJE guidance COPE similarly requires transparency about AI used in writing, images, data collection, or analysis and places responsibility on the authors. COPE position
Before submission, inspect the exact policy of the journal, conference, preprint server, data repository, funder, and institution that govern the work. Requirements differ across venues, fields, and types of contribution. Do not rely on an older policy, another journal's policy, or an assistant's summary of the rules. Preserve the version and access date of the policy you followed.
Implementation checklist
- Classify the data and task before using AI. Confirm the approved environment, consent and data-use limits, and whether a disclosure is needed.
- Start with a bounded assistive function, such as query expansion or candidate extraction, and define the human verification step before running it.
- Keep the original source, not merely a generated summary. Cite original research, documents, datasets, and methods.
- Freeze screening criteria and analysis plans before using a tool to rank records, propose codes, or suggest methods.
- Require evidence locators for all extracted fields and verify them against original full texts.
- Version-control code, preserve executable analysis scripts, test generated code, and record environments, dependencies, seeds, and input versions.
- Treat retrieved web pages, uploads, and tool output as untrusted. Do not let them change access permissions, bypass review, or create external actions without controls.
- Maintain an AI-use record with tool, model, date, prompts or settings, outputs used, validation, corrections, and disclosure decision.
- Review for bias, unsupported claims, fabricated citations, plagiarism, and changed scientific meaning before any external release.
- Recheck relevant journal, funder, sponsor, data-use, and institutional rules before submission, peer review, sharing, or a material change in tool or workflow.
Evidence
Sources used for this answer.
Question signals show what people need. Primary documentation supports the answer. Both remain visible.
- 01To Researchers, How Do You Utilize AI Tools When Conducting Research?Reddit · question signal · checked 1 Sept 2026
- 02NIH peer-review noticegrants.nih.gov · primary evidence · checked 1 Sept 2026
- 03ICMJE guidanceicmje.org · primary evidence · checked 1 Sept 2026
- 04COPE positiondoi.org · primary evidence · checked 1 Sept 2026
- 05PRISMA-Sprisma-statement.org · primary evidence · checked 1 Sept 2026
- 06PRISMA flow diagramprisma-statement.org · primary evidence · checked 1 Sept 2026
- 07NIH research-conduct guidanceoir.nih.gov · primary evidence · checked 1 Sept 2026
- 08NIH AI policy resourceosp.od.nih.gov · primary evidence · checked 1 Sept 2026
- 09NIH genomic-data noticegrants.nih.gov · primary evidence · checked 1 Sept 2026