Generative AI can produce a protocol-shaped document in minutes. That is not the same as producing a defensible protocol.
The expensive failures sit beneath the prose: an endpoint that does not answer the objective, an unsupported eligibility criterion, an infeasible visit schedule, or a citation that does not exist. If polished language hides those defects, AI has accelerated the wrong part of the work.
Use AI as a drafting assistant that transforms an approved source pack. Keep scientific decisions, source verification, privacy controls, operational feasibility, and final approval with named people.
What does human-controlled mean?
A human-controlled workflow gives the model a defined task, limits the material it can use, and requires a qualified person to verify the output before it enters the protocol.
Generative AI can confidently present false content and fabricated citations. National Institute of Standards and Technology (NIST) calls this confabulation and recommends ground-truth comparison, fact-checking, human oversight, and source review (NIST AI 600-1).
Your objective is not to make the model responsible. It cannot accept responsibility for the scientific question, participant protections, or trial conduct. Design a workflow in which errors are visible, testable, and stopped before approval.
Step 1: Define the task before opening the tool
Classify each use by what happens if the output is wrong.
| Task type | Examples | Control |
|---|---|---|
| Bounded drafting | Convert an approved synopsis into headings; rewrite for clarity | AI may draft; a named reviewer compares it with the source |
| Analytical support | Flag inconsistencies among objectives, endpoints, eligibility, visits, and analysis | AI may flag; qualified people decide whether the issue is real |
| High-consequence decision | Select endpoints, invent eligibility criteria, determine safety reporting, or choose sample-size assumptions | Do not accept model output as the decision |
| Restricted-data task | Process identifiable data, confidential agreements, or unpublished product information | Stop unless your institution authorizes the system, data flow, contract, and safeguards |
This is a Sengi recommendation, not a universal regulation. Your institution, jurisdiction, intervention, and system may require stricter controls.
Step 2: Build the source pack
A weak source pack forces the model to fill gaps. A strong pack makes gaps explicit.
Assemble the approved research question, synopsis, objectives, endpoints, eligibility rationale, intervention, assessments, visit schedule, safety pathway, statistical assumptions, current evidence, institutional templates, and unresolved decisions marked TBD.
Follow the design chain: question → design → endpoints → population → procedures → analysis → feasibility. The study-design guide provides the foundation; the protocol guide explains core document structure.
Require the model to preserve TBD, identify contradictions, and state “not supported by the supplied sources” when evidence is absent. Do not let plausible prose conceal a missing decision.
Step 3: Draft one controlled section at a time
Whole-protocol prompts make source drift harder to detect and encourage reviewers to skim polished prose. A controlled instruction names the model’s role, allowed sources, task, constraints, and required output.
For example: “Draft the objectives section using only the approved synopsis. Preserve exact endpoint names. Do not add eligibility criteria or regulatory claims. List every statement requiring investigator confirmation.”
This does not make the output correct. It makes the review surface smaller.
Step 4: Verify every claim and citation outside the model
Never ask the same model to be the final judge of its answer.
For each section, record the draft claim, exact supporting passage, reviewer and date, and disposition: accepted, revised, removed, or unresolved. Open every citation. Confirm that it exists, supports the precise claim, applies to the stated jurisdiction and intervention, and is current enough for the decision.
WHO’s guidance on large multimodal models in health identifies inaccurate or false responses, automation bias, privacy, and risks involving data entered by end users (WHO, 2024). Those risks are not solved by adding “check your work” to a prompt. They require independent verification.
Step 5: Run a protocol-coherence review
A protocol is a linked system, not a set of well-written sections.
Check whether each objective maps to an endpoint; each endpoint maps to a time point, assessment, and analysis; eligibility defines the required population; visits collect necessary data without avoidable burden; and the clinic can execute the protocol with its actual staff, equipment, patient pool, and budget.
ICH E6(R3) states that scientific objectives should be clear, protocol-execution documents should be operationally feasible, and trial systems and processes should be fit for purpose and proportionate to participant risk and data importance. It also calls for record integrity, traceability, and protection of personal information (ICH E6(R3), Principles 8.2–9.4).
AI may flag broken links in this chain. It should not decide that the chain is scientifically acceptable. For field-level scope, apply the test in Stop Collecting Nice-to-Have Data: every field should support an objective, endpoint, safety need, or planned analysis.
Step 6: Protect confidential and personal information
Do not paste sensitive trial material into a general-purpose tool because the interface is convenient.
Before restricted material enters a system, confirm approved use, contract terms, retention and training-data settings, access controls, data location, incident process, and applicable privacy requirements. De-identification alone may not resolve contractual, re-identification, or confidentiality concerns.
If those controls are unknown, use synthetic examples or remove the task from the AI workflow. This is a stop condition, not a prompt-engineering problem.
Step 7: Keep a proportionate AI-use record
Make material AI use reconstructable without turning every spelling correction into bureaucracy. Record the system and version when available, date and user, task and section, source-pack version, material output incorporated, reviewer and verification method, and unresolved limitations.
This log is a Sengi recommendation. Required documentation depends on context of use, institutional policy, jurisdiction, and whether AI output affects participant safety or study-result reliability.
FDA’s January 2025 AI document is draft, nonbinding guidance. It addresses AI used to produce information or data for regulatory decisions and explicitly excludes operational drafting or writing that does not affect patient safety, drug quality, or study-result reliability. For evidentiary use, FDA proposes controls tailored to context and risk (FDA draft guidance).
In the EU medicines context, EMA describes scrutiny that varies by context and risk and states that clinical-trial sponsors using AI/ML in medicinal-product development are responsible for ensuring relevant models, datasets, and processing pipelines are fit for purpose and aligned with applicable standards (EMA reflection paper). Do not generalize either framework beyond its jurisdiction and scope.
An IIT does not identify the legal sponsor by itself
You may control the research question and coordinate protocol development, but the individual investigator is not automatically the legal sponsor. A university, hospital, or other institution may hold that role.
Confirm the sponsor named in applicable records and agreements. Then document who reviews design, statistics, privacy, regulatory pathway, operations, and the final protocol. The role-allocation guide explains why investigator-initiated and sponsor-investigator are not interchangeable.
AI does not change that allocation. It adds another tool and another set of risks for your trial governance to control.
When should you avoid AI-assisted drafting?
Manual drafting is better when the task cannot be bounded or independently checked; evidence is incomplete; sensitive information cannot enter an approved environment; the team lacks time or expertise to verify output; system behavior or data terms are unknown; or an error could influence a high-consequence decision before review.
The strongest objection is valid: verification can cost more than drafting from scratch. The answer is not a longer prompt. Use AI only where a bounded transformation saves time after review costs are included.
Take the next step
Start with a clear research question and approved source pack—not a blank chat window. Build the design chain, assign reviewers, and use the model only where output can be checked against known evidence.
