They asked for AI.
What they needed first
was a protocol a machine
could read.
A European metabolic-disease sponsor wanted AI to improve two things: how its trial sites run, and how patients are screened. Both problems turned out to be manufactured upstream, in a conversation between the sponsor and its CRO that had no mechanism behind it. This is what we found, what we argued for, and what we built.
The five-minute version
Read this firstSix beats. Everything after this is the evidence behind them, and each beat links to where it is argued properly.
The problem
The diagnosis
The key insight
The redesigned service
The AI architecture
The value
Questions resolved in the visit; reconciliation work removed; CRA escalations falling
False screen-fails recovered; time-to-enrolment; fewer repeat visits
Feasibility known before lock; country fit visible; change orders prevented
Every answer carries its version; deviations from stale criteria fall
Each study makes the next protocol better; a shared evidence base with the CRO
AI is not the product. It is one layer in a redesigned decision system.
Five mechanisms reduce the screening error rate and only one of them needs a model. Five of the ten capabilities we built involve no machine learning at all. Four of thirteen activities were removed rather than automated, because automating them would have made the underlying defect permanent. That is not an admission that the AI is thin — it is the reason the AI is safe to put in front of a clinician, and it is the difference between a system that survives an inspection and a demo that does not.
The brief, and the two problems underneath it
Two asks that looked separate, one root cause that made both tractable at once — and a third problem nobody had written down.
Two problems, one shared cause
The sponsor came with two asks and assumed they needed two answers. They shared a root, which is the only reason the work was affordable.
Track A · what's going wrong on site
Coordinators and CRAs spend their day reconciling: which protocol version applies, who owns a question, what the CRO already knows. A question with a one-sentence answer takes three to nine days and leaves no record anywhere.
A coordination problem. Four parties, no shared view.
Track B · what's going wrong in screening
Eligibility is judged against criteria written in prose, held in five documents of unknown vintage, with evidence scattered through notes nobody has time to read. Nobody knows the error rate because nobody measures it.
An accuracy problem. One desk, one decision, no feedback.
Both fail at the same point: the protocol is authoritative as a document and meaningless as data.
A retrieval system can find the right sentence in the wrong version. A screening tool can apply a criterion superseded in March. A CRO can report against a definition the sponsor doesn't share. Fix the protocol object and all three become tractable; leave it and none of them do, however good the model is.
A third problem, which nobody had written down
Roughly three weeks into the diagnosis a pattern emerged that neither ask covered. The site problem and the screening problem were both being created upstream, at protocol design, in a conversation between the sponsor and the CRO with almost no mechanism behind it. We added it as a third track. It carries the largest prize in the engagement and it appears early in this piece, at section 08, because everything after it reads differently once you have seen it.
Three questions we had to answer
Q1
What is actually wrong with on-site trial management, and where can AI help?
03 · 09 · 11 · 12 · 17 · 19 · 20
Q2
How do you reduce the error rate in patient screening?
03 · 13 · 18 · 20 · 22
Q3
How can a sponsor and a CRO work together through AI, compliantly, in the EU and UK?
07 · 08 · 14 · 15
Eight weeks, four sites, one criterion followed all the way
We sat with coordinators during screening visits in three countries, walked the systems they use, read the support desk export, and traced a single inclusion criterion through every document that restates it. That last exercise took an afternoon and changed the shape of the engagement.
Eight working sessions with the sponsor, the CRO and two site teams, each one ending in a decision rather than a summary.
Where our own view changed
Three things we got wrong early, and when we noticed:
- Week two. We arrived expecting a retrieval-quality problem and spent the first fortnight scoping model improvements. The criterion trace in section 03 ended that line of work.
- Week four. We had France down as a capacity problem, because that is how the country team described it. Comparing it with Portugal made clear that four candidate causes were unseparated, and we withdrew the recommendation we had drafted.
- Week six. Track C did not exist in the original scope. It emerged from noticing that the same criteria kept appearing in both problem sets, and it is now the part of the work we would lead with.
What we could stand behind, and what we couldn't
Four classes, marked at the point of the claim rather than buried in a footnote.
Observed
Seen directly in a system, document or trace. Blueprints, the criterion trace, boundary classifications.
Verified
External and traceable to a published source. Regulatory dates, the screening benchmark in 18.
Modelled
Constructed to show a shape. Every percentage here. Directional, never restated as measured.
Open
Unresolved. Carried in section 24 rather than quietly filled in.
| Session | The question on the table | What came out of it | Where |
|---|---|---|---|
| 1 · Decisions | Which decisions actually matter? | Nine recurring decisions scored on evidence gap | 03 |
| 2 · The work | Where does the day break down? | Current-state blueprint for one eligibility question | 03 |
| 3 · The rule | Where does the authoritative criterion live? | The five-representation trace | 03 |
| 4 · Screening | What kind of errors are we making? | Three error types; two of them invisible | 03 |
| 5 · Beliefs | Which of our assumptions survive? | Eight hypotheses tested — three survived | 04 |
| 6 · The proposal | Would retrieval solve this? | Six objections we had to answer before it would be approved | 05 |
| 7 · Allocation | Where should the machine touch the work? | Four-stage allocation across thirteen activities | 11 |
| 8 · The line | What stays human, and why that specifically? | Checkpoints tied to consequence, not caution | 18 |
What we found
The work as it is actually done, the errors nobody can see, and which of the organisation's confident beliefs survived contact with the evidence.
On-site trial management is a reconciliation job
Q1 What's wrong on site, and where AI helpsWe expected people short of tools. We found people surrounded by tools and short of one fact: which version of the protocol is in force at this site, today.
Modelled Figures show the shape of what we observed, not audited client data.
One question, followed all the way
An ordinary event: a coordinator, mid-visit, isn't sure whether a candidate meets an inclusion criterion. Nothing here is smoothed. The phone call is drawn because the phone call is where the answer actually comes from, and no system in the organisation can see it.
Why the version question has no answer
We took one inclusion criterion and followed it into every artefact that restates it.
Protocol PDF
- Version
- Amendment 3
- Status
- Current
- Form
- Prose
Site quick guide
- Version
- Amendment 2
- Status
- Superseded
- Form
- Bullets
Training slides
- Version
- Amendment 1
- Status
- Two behind
- Form
- Paraphrase
The coordinator's sheet
- Version
- Unknown
- Status
- Last edited?
- Form
- Handwritten
Assistant corpus
- Version
- All of them
- Status
- Not modelled
- Form
- Chunks
The data underneath, scored on fitness rather than completeness
A field can be fully populated and worthless. The clearest example in this case looks healthy in every report the sponsor produces.
| Asset | Score /5 | Complete? | Timely? | Fit for the decision it feeds? |
|---|---|---|---|---|
| Protocol content as data | 1.1 | n/a — prose | — | No. Not represented as data at all, so no rule, gate or comparison can be built on it. |
| Screen-failure reason | 1.8 | Largely populated | Coded up to 42 days later | No. Reconstructed after the fact by someone who wasn't in the room. It's a plausible guess, not the reason. |
| Site performance history | 2.1 | Partial | Monthly | Barely. Enough to rank sites. Not enough to explain one. |
| CRO operational record | — | Held by the CRO | Monthly, as PDF | No. Arrives as a document after the window in which it could change a decision. |
| Support interactions | — | Only what reaches the desk | Real time when logged | Partial. The phone-resolved majority never enters it. |
| Screening counts | — | Three versions exist | Varies by source | No. Three numbers, three definitions, no owner of the definition. |
Which decisions actually matter, and what they were missing
We anchored the whole design on decisions rather than processes, because a decision is the only unit that tells you what evidence a decision requires. Three questions per decision: what evidence would a competent person need, does it exist anywhere, and can they reach it in time. The third is where nearly everything fails, and conventional data-maturity assessments don't ask it.
What this does to the person holding it
We ran the coordinator's role through a work-system analysis — person, tasks, tools, environment and organisation examined together rather than blaming any one of them. The finding is specific and uncomfortable.
Person
- Trained on the protocol at initiation, once
- Running three to six studies at a time
- No notification when any of them changes
Tasks
- Screen, consent, schedule, query
- Chase, reconcile, re-enter, translate
- Five of eight are pure overhead
Tools
- EDC, eTMF, CTMS, email, phone
- Paper binder, personal spreadsheet
- And now an assistant in a new tab
Environment
- Clinic-floor interruptions
- The question arises with the patient in the room
- Answer arrives after they've gone home
What the systems carry
- Storing the document
- Delivering the email
- Counting the logins
What the human carries
- Remembering which version is current
- Remembering who answered this last time
- Spotting a change nobody announced
- Reconciling three counts that disagree
- Apologising to the patient
Automation bias, created not avoided
A version-unaware assistant that is usually right teaches people to stop checking. The speed arrives before the failure mode does, which is the worst possible ordering.
Trust calibration, absent
The product showed no evidence, no version and no confidence state, so users could only fully trust it or ignore it. They chose to ignore it.
Situation awareness, degraded
An answer with no source strips the context the investigator needs for the judgement they remain accountable for. Speed bought from the wrong account.
So where can AI help on Track A? Q1
In three places, and only after one non-AI thing is fixed.
Fix first
Make the protocol a versioned data object. No model involved. A data-modelling job that unblocks everything else.
Then
Answer against the applicable version, with the source shown. Removes the reconciliation, not just the search.
Then
Route to a named owner automatically, and show that owner four sites asked the same thing.
Then
Map amendment impact before the change lands — sites, documents, training, consent.
Two of the three screening errors are invisible
Q2 Reducing the screening error rateYou cannot reduce an error rate nobody measures. Before designing anything, we had to establish what kinds of error screening produces and which of them the organisation can currently see.
Step one of reducing the error rate isn’t technology at all.
It's a re-adjudication study: take a sample of past screen failures, have a blinded clinician re-review them against the criteria that were in force on the day, and establish a baseline. Without that number, any later claim of improvement is unfalsifiable — and it is exactly the number the previous pilot never produced.
Where screening loses people
Modelled Widths show the shape observed, not audited conversion figures.
The constraint that shapes everything here
The sponsor has no lawful route to identifiable records for anyone screened and never enrolled. That is not a gap to close. It is the shape of the problem.
Screening intelligence cannot be built by moving patient data to the sponsor, because the patients who matter most — the ones turned away — are the ones whose data can never move. So the computation goes to the site and only criterion-level outcomes come back. Section 15 is what that looks like.
What the organisation believed, and what survived
The brief arrived with confident positions. We treated each as a hypothesis and tested it against what the material could actually support. Three survived, three were reframed, two were contradicted. The reframings mattered most, because they redirected effort that was about to be spent in the wrong place.
“The pilot worked.”
“The problem is retrieval quality.”
“What we need is contextual reasoning.”
“Patient screening is the biggest value pool.”
“France underperforms because of capacity.”
“CRO execution is the main problem.”
“CTIS is a compliance workstream.”
“More automation is the answer.”
Making the case
A retrieval assistant in a GxP environment is a governance argument before it is a technical one. Six objections had to be answered before anything would be approved, and answering them is what turned a recommendation into a design.
Arguing for retrieval in a room that had reasons to say no
Recommending a retrieval assistant into a GxP environment is not a technical argument, it is a governance one. We knew which objections would land before we walked in, because they are the same six every time, and we built the case around answering them rather than around the technology.
The list below is what a sponsor's quality, medical, data and privacy functions each need satisfied before anything like this gets approved. Working through them is what turned a technology recommendation into a service design, and each one produced a design constraint we would not otherwise have written.
Clinical quality
A citation from a superseded amendment is a GCP finding, not a UX bug.
An answer that looks right and comes from the wrong version will be acted on, and it will look correct in the audit trail afterwards. This is the objection that has to be answered first, because nothing else matters if it stands.
What we broughtThe five-representation trace. It demonstrated the failure already happening without AI, which reframed the assistant as a way to close an existing hole rather than open a new one.
Study medical
A large share of site questions have no answer in the protocol at all.
A system that always produces something will produce something for those too, in the register and tone of the study medical lead. The concern is not hallucination in the abstract, it is fluent false authority on questions that require a person.
What we broughtA read of the support desk export showing how many questions were interpretation rather than lookup, and a commitment to measure abstention in both directions.
QA and computer systems validation
If this is a computerised system under GCP, where is the validation package?
And what happens the week the model provider ships an update. A system whose behaviour can change without a release cannot be validated in the ordinary way, and that is a real objection rather than an obstructive one.
What we broughtVersion pinning, a fixed adversarial evaluation set, and the rule that a protocol amendment re-enters evaluation exactly as a model change does. That last point is what carried it.
Country study management
When an inspector asks why a site did something, someone has to sign the answer.
Accountability cannot sit with a tool. If a coordinator acted on machine output, the record needs to show who was responsible for that guidance and what evidence they had in front of them.
What we broughtAn owner directory mapped to question type, and an answer record that carries its source version. Accountability became locatable rather than implied.
Clinical data
Someone is hand-building the corpus for two studies. Who does that for forty-one?
Manual curation is the scaling ceiling, and it was invisible in the pilot because two studies is a manageable amount of work for one motivated person. At portfolio scale it is a standing headcount with no owner.
What we broughtAn estimate of the implied curation effort across the portfolio, which moved the conversation from “should we scale the assistant” to “what has to change before it can scale”.
Data protection
The CRO holds half of what this thing wants to reason over.
Controller and processor roles have to be settled per processing activity, at design time. Discovering the question at go-live means either a delay or a system that quietly operates outside its basis.
What we broughtA map of every data flow between sponsor, CRO, site and model provider, drawn before any build started, and the finding that the CRO blocker was contractual rather than legal.
Every objection came back to the same thing: the system had no way to know what was true, when, for whom, and no way to say so when it didn't.
Six objections, six requirements
We stopped defending the recommendation and rewrote it against what they'd said. This is the pivot of the engagement.
| The objection | The requirement it implies | How the design meets it | Where |
|---|---|---|---|
| Wrong amendment = GCP finding | Applicability resolved before retrieval, deterministically, and visible in the answer | Version resolution is a hard gate ahead of the index, not a filter behind it. Non-applicable content is unreachable rather than down-ranked. | 12 |
| It'll invent an answer | The system must be able to decline, and declining must be measured as carefully as answering | Four evidence states drive permitted behaviour. Where sources conflict, generation doesn't run — both passages are shown and the case is routed. | 15 |
| Where's the validation package? | Fixed adversarial evaluation set; version correctness as a release gate; re-validation on any change | An eight-measure harness. A protocol amendment is treated as a model change and re-enters evaluation. | 19 |
| Who signs the answer? | Named accountability per question type, and a record of the evidence the human saw | Owner directory plus deterministic routing. Every answer writes a record carrying its source version; every override records a reason. | 14 |
| Who curates 41 studies? | Structuring the source automated, owned by clinical development, validated field by field | Criterion extraction into a schema, with anything failing validation queued for a human rather than accepted. | 13 |
| Controller or joint controller? | Roles, lawful basis, transfer route and data rights settled at design time and written into contract | Sponsor–CRO data contract on five event types; patient computation kept site-side so the hardest question never arises. | 17 |
Six solution archetypes, assessed
Q3 Working with the CRO, compliantlySix archetypes, assessed as at June 2024. Named examples illustrate each category — they are not endorsements, evaluations or a competitive assessment.
| Archetype | Examples | What it solves | Where it stops | Track | Call |
|---|---|---|---|---|---|
| Clinical content & trial platforms | Veeva Vault; Oracle Clinical One | Custody of documents, milestones and study records. Already in place and doing its job. | Stores the protocol as a document. No representation of a criterion, and no answer to “which version applies here today”. | A | Already own |
| General enterprise RAG & copilots | Microsoft Copilot with Azure AI Search; Glean | Fast, competent semantic search across a document estate, with solid enterprise plumbing. | No concept of regulatory authority hierarchy, effective dates or country applicability. Will confidently return the training deck. | A | Not the answer |
| Protocol digitisation & authoring AI | Faro Health; Certara | Structuring protocol design and generating regulatory documents from structured inputs. | Aimed at authoring new protocols. Doesn't retro-structure a portfolio of forty-one live studies, which is the actual job. | A + C | Partner, partial |
| EHR-based screening & matching | Deep 6 AI; TriNetX; Mendel | Finding candidates in unstructured clinical records at scale. Strong — this category works. | Mostly built around US health-system data access. An EU multi-country site estate and the never-enrolled boundary change the deployment model entirely. | B | Partner, site-side |
| Operational analytics & RBQM | Medidata Detect; CluePoints; Saama | Statistical signal detection across trial data — site outliers, data anomalies, risk indicators. | Detects that a site is an outlier. Can't say why, because the operational events that would explain it don't exist in the sponsor's estate. | A | Buy later |
| Foundation models, built on directly | Anthropic Claude; OpenAI | Extraction, classification, grounded composition and explanation — the reasoning layer, at low cost per call. | Ships no domain structure at all. Everything that makes it safe here has to be built around it. | A + B + C | Build the thin layer |
The gap in the market, stated plainly
No vendor offers a protocol version authority service — a component that takes a study, a country, a site and a date and returns exactly one applicable criterion set, with provenance. Every vendor assumes you already know which document is in force, because in their model you hand them the document.
Nor does anyone offer what Track C needs: a way for a sponsor and a CRO to test a draft protocol against real site populations before it is locked, without either party seeing the other's data.
Both are small. The first is roughly two hundred lines of deterministic logic over a data model that didn't exist. That's what we built; everything else was bought, partnered or deferred.
Buy
Document custody, statistical monitoring, model inference. Mature markets, real products, no advantage in building.
Partner
Site-side screening deployment and protocol structuring. Specialist capability, but integration and governance stay with the sponsor.
Build
Version resolution. The authority-tier scoring. The operational event model. The owner directory. The pre-lock feasibility loop. Small, unglamorous, and nobody else can supply them.
The largest prize
Neither ask covered this. Both problems are manufactured upstream, at protocol design, in a conversation between sponsor and CRO that has almost no mechanism behind it.
Sponsor and CRO, designing the trial together
Q3 Working with the CRO, compliantlyNeither ask covered this. But both problems are manufactured upstream, at protocol design, in a conversation that has almost no mechanism behind it. This is the section with the largest prize and the least prior art.
How they actually communicate
The sponsor writes the protocol. The CRO and the sites execute it. Between those two facts sits a set of channels that are all either slow, low-bandwidth, or leave no record — and usually all three.
The protocol is written by people who cannot see the populations, and executed by people who can — and there is no moment where those two facts meet before the design is locked.
By the time reality arrives it arrives as a change order. The sponsor experiences it as CRO underperformance; the CRO experiences it as an undeliverable protocol. Both are describing the same missing mechanism, and the commercial relationship makes it costly for either to say so first.
What we built instead: run the protocol before you lock it
Two things we'd already built made this possible without either party seeing the other's data. Criteria are now structured objects (14), and there is a local engine at each site that can evaluate a criterion against that site's own population and return only counts above a cohort threshold (15). Put those together and a draft protocol becomes a query you can run.
Draft criteria · 38 sites, 3 countries
IC1. Raising it protects safety and shrinks the pool.
IC4. Shorter windows mean fewer patients have a recent enough result on file.
EC7. A safety exclusion with a direct recruitment cost.
EC2. Longer washout excludes more, and is harder to evidence from notes.
Eligible per 1,000, by country — the same criteria, different local practice
Modelled Pass rates and country modifiers are illustrative. In the real loop each number is a live federated count returned by the sites themselves.
The second shared object: what a change actually costs
The other conversation that goes badly is amendments. The sponsor proposes a change; the CRO estimates effort; neither has a shared picture of the blast radius, so the negotiation is about the number rather than the change. Once the protocol is a graph, the blast radius is computable and both parties look at the same figure.
Proposed change
Today · manual tracing
With the protocol graph
Modelled Counts illustrate relative blast radius across change types.
The operating model that goes with it
Software on its own does not change how two organisations talk. What changes the conversation is having a shared object to point at, and a defined moment to point at it.
| Moment | Before | After | Shared object |
|---|---|---|---|
| Protocol drafting | Sponsor writes. CRO comments on operability in prose, late, and diplomatically. | Each criterion draft is run as a federated query. The CRO's operational judgement arrives attached to a number. | Feasibility run, versioned per draft |
| Criterion review | A meeting where nobody can evidence a claim. | A joint session over the country breakdown, with the criteria that cost most recruitment ranked. | Criterion cost table |
| Lock | Sponsor decides. CRO accepts commercially. | Sponsor still decides — but the trade-off is recorded, with what each party advised. | Decision record with evidence refs |
| Execution | Monthly PDF, reconstructed into a tracker by hand. | Five operational event types on an agreed cadence, into a schema. | Operational event stream |
| Issue review | Volume-based status reporting. | Recurrence-based. “This criterion generated 34 questions across 14 sites” is a protocol defect, not a CRO failure. | Question volume by criterion |
| Amendment | Change order and a negotiation about effort. | Both parties open the same blast radius before anyone prices it. | Impact map with coverage stated |
| Close-out | Lessons learned document nobody reads. | The feasibility runs and the actual outcomes are compared, and the deltas feed the next protocol. | Prediction-vs-actual record |
What the sponsor gains
A protocol that recruits. Amendment risk visible before lock rather than eighteen months after it. And a factual basis for the CRO conversation that doesn't depend on trusting a PDF.
What the CRO gains
A channel for the thing its CRAs already know, with the operational reality attributable to the protocol rather than to their delivery. Fewer undeliverable studies. And a defensible position when a criterion is the problem.
What neither gives up
No patient data crosses. No commercially sensitive CRO data crosses. The sponsor's protocol authority is untouched. The only thing that moves is a count, and the only thing that changes is when the conversation happens.
The design
Where the machine touches the work, which techniques were assessed and rejected, how retrieval is scored across three sources of truth, and how any of it scales to forty-one studies.
The AI service blueprint
Two tracks drawn separately. They are different services with different users, data and legal exposure. What they share is the layer underneath. Track C gets its own section, because it changes who is in the room rather than what happens in it.
One eligibility question, before and after
The same event, drawn twice. The test we applied throughout was not what the product adds, but what the coordinator stops doing.
Track A · on-site trial management
The same event as Fig 01, redesigned. The test we applied wasn't “what does the product add?” but “what does the coordinator stop doing?”
Track B · patient screening
A different service. Computation runs at the site because the data can't leave it, and the output is criterion-level evidence rather than a verdict.
Fourteen techniques considered, five built
Before designing anything we assessed the available AI techniques against this organisation rather than against the field. Two questions per technique: could it do the job, and could this company run it. The second killed more candidates than the first.
| Technique | What it would do here | Data prerequisite | Maturity | Scales across 41 studies? | Call |
|---|---|---|---|---|---|
| Retrieval-augmented generation | Answer protocol questions against the applicable version | A versioned, structured corpus | Mature | Yes, once the corpus is structured | Build |
| Structured extraction | Turn protocol prose into criterion objects | The protocol documents, and a schema | Mature | Yes — batch, once per amendment | Build |
| Entity resolution | Reconcile the same site and investigator across eTMF, CTMS and EDC | Identifier overlap | Mature | Yes | Build — prerequisite |
| Knowledge graph | Amendment impact, dependency traversal, criterion lineage | The entity model | Mature | Yes — ten node types, not an enterprise graph | Build, narrow |
| Federated analytics | Criterion pass rates across sites; pre-lock feasibility | A local engine per site and agreed query patterns | Emerging in this setting | Yes — the governance scales linearly, not the compute | Build — Track C |
| Bounded agents | Coordinate the consequences of an amendment | The graph, plus a read-only tool layer | Emerging | Yes, in one narrow workflow | Build, one only |
| Document intelligence / OCR | Bring scanned historic protocols and site binders into scope | The scanned corpus | Mature | Partly — quality varies by site | Buy |
| Classical anomaly detection | Flag site behaviour that departs from peers | Operational event history that does not yet exist | Mature | Yes, later | Defer to step 6 |
| Time-series forecasting | Project enrolment per site and per country | Clean enrolment history with consistent definitions | Mature | Yes, later | Defer to step 6 |
| Causal inference | Establish why France differs from Portugal | Event history plus site covariates | Mature, but demanding | Medium — needs enough sites per stratum | Defer, then narrow |
| Federated learning | Predict candidate fit across heterogeneous site populations | Validated labels, comparable semantics, a shared evaluation design | Research-grade in clinical deployment | Not yet — see 14 | Defer |
| Speech-to-text | Capture monitoring visits and site calls as structured records | Audio capture, and consent from everyone on the call | Mature | Low — the consent burden scales badly across three jurisdictions | Declined |
| NLP for adverse-event coding | Support MedDRA coding in the safety database | Safety data, which sits outside this scope | Mature | n/a | Out of scope |
| Scheduling optimisation | Plan monitoring visits by risk and travel cost | Risk model plus travel data | Mature | Yes, but the value here is small | Declined |
Where each one sits
Could this organisation actually run it?
Capability assessments usually stop at the technique. The more useful question is organisational: five dimensions, scored one to five, per capability area. A capability with strong data and weak governance fails just as completely as one with no data.
Where the machine was allowed to touch the work
Q1 What's wrong on site, and where AI helpsAutomation is not a property of a task. It's a property of a stage within a task. Every significant activity was split into four stages and allocated separately at each one. This is the piece of work that stops “AI-assisted” from being a phrase that means nothing.
Four of thirteen activities were removed rather than automated.
The corpus curation, the criteria sheet, the count reconciliation and the ownership chase all exist because something upstream is missing. Automating any of them would have bought a small efficiency and made the underlying defect permanent and harder to see.
Exactly what the AI does — and what it doesn't
Ten capabilities. Five use a language model. Five are ordinary deterministic software that gets called AI far too often. Naming which is which is the difference between a system you can validate and one you can only hope about.
Version resolution
Which criteria are in force here, today
- How
- Deterministic lookup over the version object. Study + country + site + date → exactly one criterion set, or a declared conflict.
- Model
- None. ~200 lines of code.
- Human
- Clinical development owns the version data.
Criterion extraction
Protocol prose → structured data
- How
- Long-context model with a strict output schema: criterion id, logic, threshold, unit, window, evidence required.
- Model
- Claude Opus, batch, once per amendment.
- Human
- Clinical development approves each extracted set.
Query understanding
Turning a typed question into filters
- How
- Classify question type; extract study, criterion, drug, timepoint. These populate the hard filters before retrieval.
- Model
- Claude Haiku. Thousands of calls a day, negligible cost.
- Human
- None; low-confidence labels route to a default owner.
Scoped hybrid retrieval
Finding the passage, inside one version
- How
- BM25 for exact identifiers, thresholds and units; dense vectors for paraphrase; fused, then reranked on the domain features in 12.
- Model
- Embedding + cross-encoder. No generative model.
- Human
- None.
Grounded composition
Writing the answer, only from evidence
- How
- Answers only from retrieved passages, low temperature, with citation binding — any sentence without a source span is dropped before display.
- Model
- Claude Sonnet, prompt-cached on the protocol context.
- Human
- Reader confirms applicability to their patient.
Criterion-level screening
Reading notes against one criterion at a time
- How
- One targeted question per criterion against authorised clinical notes. Returns pass, fail or insufficient evidence, with the source note and its date.
- Model
- Runs inside the site environment. Never returns an eligibility verdict.
- Human
- Investigator makes the clinical determination.
Routing
Getting the question to its owner
- How
- Rules over a published question taxonomy and an owner directory. A lookup, and nothing more.
- Model
- None. A model here adds latency, cost and a failure mode for no gain.
- Human
- The named owner resolves and closes.
Amendment diff
What actually changed between versions
- How
- Deterministic structural diff over criterion objects. A model then explains materiality — it never determines what changed.
- Model
- Sonnet for the explanation only, shown alongside the raw diff.
- Human
- Clinical development judges materiality.
Federated feasibility query
Testing a draft protocol before it's locked
- How
- A structured criterion set is sent to each site's local engine; each returns counts above a cohort threshold. No patient data moves in either direction.
- Model
- None for the count. The criterion set that drives it comes from capability 02.
- Human
- Sponsor and CRO jointly interpret; the protocol call stays with clinical development.
Amendment impact agent
The only agentic capability in the design
- How
- Triggered by a change event. Traverses criterion → site → training → consent, and drafts a review task per affected item.
- Model
- Claude with read-only tools; writes are drafts only.
- Human
- A named owner approves the impact map before anything is published.
Three sources of truth, and the rules between them
Q2 Reducing the screening error rateMost descriptions of RAG assume one corpus. There are three here, they are not comparable, and the interesting engineering is in how a question gets routed between them and what happens when they disagree.
| Substrate | Holds | Answers questions like | Retrieval method | Authority |
|---|---|---|---|---|
| Structured criterion store | Criterion objects: threshold, unit, window, logic, effective date, applicability | “What is the HbA1c window?” “Does a value from 9 weeks ago qualify?” | Database lookup or rule evaluation. No model, no embedding, no retrieval. | 1.0 — by construction |
| Document corpus | Protocol body, amendments, country annexes, investigator brochure, lab manual | “Why was the window changed?” “What does the protocol say about repeat testing?” | Hybrid BM25 and dense, scoped to the applicable version, then reranked | Tiered — see the weights below |
| Precedent store | Prior approved answers, the event log, resolved conflicts, override reasons | “Has anyone asked this?” “How was it answered for the other French sites?” | Semantic match on the question, filtered to answers whose source version is still in force | Medium, and it decays as versions move |
Routing happens before scoring
The classifier's job is not to pick an answer. It decides which substrates can legitimately answer the question, and a large share of questions never reach retrieval at all because the structured store holds an exact value.
Study, country, site, role, date — taken from the session, never typed by the user
Question type and entities extracted. Sets the hard filters and selects the routes.
Exact value returned directly. No retrieval, no generation. Roughly a third of site questions end here, and they end in under a second.
Version-scoped hybrid search, reranked on the five domain features. Used for rationale, edge cases and anything the criterion object cannot express.
Prior approved answers whose source version is still in force. Returns the answer and who approved it, which matters more than the text.
Results fused on rank rather than score, because a database exact match, a BM25 score and a cosine distance are not on one scale. Precedence rules then apply.
Which substrate answered, which version, which passage, and the evidence state
Three precedence rules, none of them weights
01 Structured beats prose
If the criterion store holds a deterministic value, that value is the answer. The document corpus may add rationale alongside it. It may never override it.
Why not a weight: the same reason version applicability isn't one. Anything expressed as a weight can be outweighed by a sufficiently confident competitor, and the failure is silent.
02 Precedent informs, never decides
A prior approved answer is shown as precedent with its approver and date. It is not merged into the answer text, because an approved answer to a similar question is not an approved answer to this one.
Why: precedent silently becoming policy is how a one-off interpretation turns into an unwritten rule nobody remembers agreeing to.
03 Disagreement is an alarm
If the corpus says twelve weeks and the criterion store says eight, that is not a ranking problem to be resolved. It is an extraction defect or a stale amendment, and it pages the corpus owner.
What this buys: the retrieval system becomes a continuous check on the quality of the organisation's own data. It surfaces protocol defects as a side effect of answering questions.
Inside route B: the version gate runs first
Study, country, site, role, date
Boolean. One applicable criterion set, or a declared conflict and a full stop.
BM25 plus dense, fused with reciprocal rank fusion. Only the applicable partition is reachable.
Cross-encoder relevance adjusted by five clinical features
Every claim tied to a passage. Unbound claims dropped.
Answer, source, version, evidence state, route
The scoring function, and why it isn't off the shelf
Relevance alone ranks a training slide above a protocol section, because the slide is written in the same plain language as the question. Five domain features fix that.
Question from a French site, 17 June 2024
“Can we include a patient whose HbA1c was taken 9 weeks ago?”
Cross-encoder semantic match. The only feature generic RAG has.
Protocol > country annex > brochure > training deck > site-authored.
Contains the exact criterion ID, threshold, unit or window.
How recently this source took effect.
Source language against site language; translations score below the controlled source.
A human-approved answer already issued, whose source version is still in force.
Modelled Illustrative feature values on six representative passages, to show the mechanism.
Two numbers decide whether it answers at all
Absolute support
Does any passage clear a minimum support threshold? If not, the system abstains. This catches the questions the protocol does not address, which is a large share of what sites actually ask.
Top-1 margin across authority tiers
If the best and second-best passages come from different authority tiers and their scores are close, that is a conflict signal rather than a ranking problem. It triggers the conflict state instead of a confident answer.
How this scales to forty-one studies without moving any patient data
Q3 Working with the CRO, compliantlyFederated architecture is the backbone of both Track B and Track C, so it belongs in the open rather than in an appendix. What follows is the ladder, what actually blocks each rung, and where we stopped.
The topology
A draft or in-force protocol expressed as machine-readable criteria. This is all that travels down.
Pre-approved query patterns only. Every request logged, and the log is auditable by both sponsor and CRO.
France · 14 sites
Local engine at each site. Criteria evaluated against local records; nothing transmitted outward.
Portugal · 9 sites
Same engine, same criterion set, different local terminology mapping.
Spain · 15 sites
Two hospital groups with shared IT; the rest independent, which is where integration cost lands.
Patient records, notes, labs, medication history and any identifier stay inside the treating organisation. Not de-identified and forwarded — not transmitted at all.
criterion_id · outcome · protocol version used · timestamp · count. A site with fewer than the threshold returns “suppressed”, never a small number.
The five levels
Federated rules
Deterministic criteria pushed to each site and evaluated locally. Age, lab values, dates, windows.
Federated retrieval
Local RAG over authorised clinical notes, finding the evidence for each criterion. Inference runs inside the site.
Federated analytics
Aggregate queries across sites. Which criteria disqualify most candidates, and whether a draft protocol would recruit.
Federated learning
Train a shared model across sites without pooling data. Gradients or weight deltas move, not records.
Federated evaluation
Run the test set at every site and report performance per site rather than in aggregate. The level almost nobody builds.
What actually blocks federated learning here
Not compute, and not model architecture. Six things, in the order they bite.
Blocker 1
Semantic heterogeneity
The same criterion resolves against local lab codes, different coding systems, different units and different note conventions at every site. “eGFR” is not one field across thirty-eight hospitals.
What it costsA terminology mapping layer per site. This is the real price of scaling, and it is paid before any model exists.
Blocker 2
Non-IID populations
Site populations differ systematically — referral patterns, comorbidity mix, socioeconomics. A globally averaged model can be worse at every individual site than that site's own local model.
MitigationPersonalisation layers or cluster-based aggregation. Both need level 5 to prove they helped.
Blocker 3
N sites means N approvals
The compute cost of adding a site is negligible. The governance cost is a data processing agreement, an IT review, an ethics position and a local sign-off. That scales linearly and never gets cheaper.
What we didKept the returned payload identical and minimal across all levels, so one approval template covers every query pattern.
Blocker 4
Systems heterogeneity
An academic centre and a district hospital are not comparable nodes. Synchronous aggregation means the slowest site sets the pace for everyone.
MitigationA minimum viable node specification, and asynchronous aggregation so a slow site degrades its own contribution rather than the run.
Blocker 5
Gradient leakage
Model updates carry information about the data that produced them. “The data never leaves” is true and insufficient.
MitigationSecure aggregation plus differential privacy, with the privacy budget stated as a number in the protocol rather than asserted as a property.
Blocker 6
Combinatorial versioning
Model version times protocol version times site. Without a registry you cannot answer which model produced a given answer under which protocol at which site — which is the first question an inspector asks.
What we builtThe registry, at level 1, long before it was needed. It is cheap early and impossible to retrofit.
The economics of stopping at level 3
Level 4 becomes justifiable in one specific circumstance: when you need a prediction that a count cannot give, and no single site has enough positive examples to build it alone. The realistic candidate here is pre-screening prioritisation — ranking which chart to review first. That is a genuine federated-learning problem, and it is worth revisiting once levels 1 to 3 have run long enough to produce labels. It was not worth starting with.
Federated learning was sequenced, not dismissed. Federated architecture was built on day one.
Conflating the two is the most common mistake in this area. The distributed rules and the distributed counts are what make Track B lawful and Track C possible, and neither of them trains anything.
Running it on Claude
Q3 Working with the CRO, compliantlyClaude is not the retriever and not the system of record. It does five bounded jobs at three cost tiers, wrapped in deterministic code that decides what it's allowed to see.
| Job | Tier | Why this tier | Pattern | Volume |
|---|---|---|---|---|
| Query understanding | Haiku | Classification into a small, stable label set. Latency matters; capability doesn't. | Structured output, single call | Very high, per question |
| Grounded composition | Sonnet | Must follow a strict answer contract and stay inside supplied passages. Low temperature. | Prompt caching on the protocol context | High, per answer |
| Criterion-level screening | Sonnet | One narrow question per criterion beats one broad question about eligibility — more accurate, and each answer is separately checkable. | Site-side; parallel per criterion | Medium, per candidate |
| Diff explanation | Sonnet | Explaining a computed diff is a language task. Computing it is not, and isn't given to the model. | Diff supplied as input, shown alongside | Low, per amendment |
| Criterion extraction | Opus | Long documents, subtle logic, high cost of a silent error. Worth the best model available. | Batch API, tool schema, strict validation | Low, per amendment |
| Amendment impact | Opus | Multi-step relational reasoning across the protocol graph with read-only tools. | Bounded agent, draft writes only | Low, per amendment |
| Evaluation judging | Sonnet | Scores groundedness and citation validity offline, with human adjudication on a sample. | Offline only — never judges live output | Per release |
The model never sees inapplicable content
Version scoping happens before the prompt is assembled. Superseded passages aren't in the context window at all, so there's no possibility of the model reasoning its way to them.
Deployment follows the data
Track A runs in the sponsor's EU environment. Track B runs inside the site's. The same capability deployed twice, because the constraint is legal rather than technical.
Model changes are change control
Version pinning, a fixed evaluation set, and re-validation before any tier is upgraded. The quality team's objection about the model changing underneath them is a release process, not a reassurance.
In the clinic
How people get onto the system, what each track looks like in daily use, how trouble is detected, and the legal instrument on every line.
Getting staff and patients onto the system
Q1 What's wrong on site, and where AI helpsAccess control in clinical research is usually rebuilt from scratch for every tool, badly. There is an artefact that already exists, is already maintained, and is already inspected, and almost nobody wires systems to it.
Staff: the delegation log is the access layer
Good clinical practice already requires a delegation of authority log for every study, naming who may perform which trial-related duties and from when. It is signed by the investigator, kept current, and examined at inspection. It is also, in most organisations, a spreadsheet that no system reads.
Feasibility complete, site chosen for the study
Clinical trial agreement and data processing terms signed
Investigator names each person and the duties delegated to them, with dates
Protocol and GCP training completed and recorded per person
Role and scope derived from the delegation log. Access is limited to that study and those duties.
Delegation log entry ends, access ends the same day. No leaver process to forget.
What this fixes
Access that outlives the person's involvement. In the pilot, provisioning was a spreadsheet and a support ticket, and nobody owned deprovisioning at all.
What it costs
The delegation log has to become structured rather than a scanned PDF. Roughly the same work as structuring a protocol, and the same argument applies.
Where it strains
Sites that maintain the log on paper. Around a third of the estate, and they need a lightweight entry route rather than an exception process.
Patients: mostly they never touch it, and that is a design decision
Everything in Track B runs behind the site's own systems and is used by the coordinator and investigator. The patient experiences it as a shorter wait and fewer repeat visits. That is the intended relationship, and there is a reason to be careful about changing it.
- Screening resolved within the visit rather than across two
- Fewer repeat tests ordered because a criterion was misread
- Fewer people turned away who were in fact eligible
- No new account, no new app, no additional consent
Let candidates check likely eligibility themselves before a visit. It would widen the funnel at the top, which is the part nobody currently instruments.
- Reduces wasted visits for both sides
- Reaches people no site would have identified
- Creates the first real data on who was considered
A patient using a pre-screen becomes a data subject of whoever hosts it. If the sponsor hosts it, the sponsor becomes a controller for people who never enrol — the exact boundary the whole architecture exists to respect.
So it is hosted by the site, under the site's own basis, with the same governed output contract: counts up, nothing else. Same pattern, applied one layer further out.
On-site trial management
Q1 What's wrong on site, and where AI helpsFour changes. Only one of them is a model. Together they close the reconciliation loop that was consuming the coordinator's day.
01 The protocol becomes a data object
Each criterion is a record with an identifier, structured logic, threshold, unit, evidence requirement, effective date and applicability scope. Representations — training deck, quick guide, site sheet — become derived artefacts pointing at the criterion, not independent documents that drift.
Closes: the five-representation problem. Change a criterion and everything pointing at it is flagged automatically.
02 Every question type gets an owner
A published taxonomy of question types, each mapped to exactly one accountable function, per country. Routing is a lookup. The coordinator stops choosing who to phone, and the answer stops depending on who she happens to know.
Closes: the France–Portugal gap. Portugal already ran one named contact per site and performed better; this generalises an existing practice rather than inventing something.
03 Answers carry their evidence
Version, effective date, the passage, the escalation route and the feedback control — the same shape every time, whether the system is confident or not. Users learn the shape, then learn which part to check.
Closes: trust calibration. The previous assistant exposed no source, so users could only fully trust it or ignore it. They ignored it.
04 The work writes its own record
Seven events, created as a by-product of doing the job. Nothing is a data-entry task. The gate that changed
behaviour: AnswerIssued requires a source_version, so a verbally-resolved answer
can no longer enter the record at all.
Closes: the invisibility of the phone call — and with it the CRO conversation, the France analysis and the evaluation set, none of which are separate builds.
study_id, phase, therapeutic area, milestones
protocol_id, therapeutic intent
effective_from, supersedes, applicability scope
human_text, structured_rule, threshold, unit, window, evidence_requirement
source_doc, section, page, translation, derived_from
regulatory context, annexes in force
site_id, organisation, capability, contract
role, affiliation, study history
These three already existed. They were given identifiers and owners rather than rebuilt.
amendment_id, old, new, rationale, affected_criteria[] — the object that makes impact analysis possible at all
question, issue, milestone, actor, timestamp, source
owner, evidence_used[], override, outcome
Six hops. That single path answers “which sites does this amendment affect, which documents go stale, who needs retraining, and who has already asked about it” — the question that previously took three weeks of manual tracing and found two-thirds of the answer.
The canonical model and the seven-event contract engineering detail
No lake. The architecture starts from the smallest set of objects that make the nine decisions in 03 computable, and every object has exactly one owner. Split ownership of a fact is where programmes like this quietly rot.
| Object | Key fields | Owner | New? |
|---|---|---|---|
| Study | study_id, phase, therapeutic area, status, milestones | Clinical systems | Existing — given identifiers |
| Protocol | protocol_id, therapeutic intent | Clinical development | Existing |
| Version | version_id, effective_from, supersedes, applicability scope | Clinical development | New |
| Criterion | human_text, structured_rule, threshold, unit, window, evidence_requirement | Clinical development | New |
| Representation | source_doc, section, page, translation, derived_from | Derived — never authored | New |
| Change | amendment_id, old, new, rationale, affected_criteria[] | Clinical development | New — makes impact analysis possible at all |
| Country / Site / Investigator | regulatory context, annexes in force, capability, contract, study history | Regulatory / clinical ops | Existing — given owners |
| Event | question, issue, milestone, actor, timestamp, source | Operational functions | New |
| Decision | owner, evidence_used[], override, outcome | The decision owner | New |
| AI interaction | capability, model_version, output, feedback | AI governance | New |
| Event | Required fields | Rejected if |
|---|---|---|
ProtocolChanged | study_id, old_version, new_version, effective_date, affected_criteria[] | No named owner or no source document |
QuestionRaised | study_id, site_id, role, category, timestamp, content | Identity doesn't resolve; category isn't in the taxonomy |
AnswerIssued | question_id, source_version, answer, responder, timestamp | source_version missing. An answer without a version is not a valid answer. |
IssueResolved | issue_id, action, owner, outcome | Outcome is null — an unresolved issue stays open |
SiteMilestone | site_id, milestone, date, source | Milestone name isn't from the controlled list |
EligibilityAssessment | local_subject_token, study_id, criterion results[], evidence timestamps | Token isn't site-local — central resolution is rejected at the boundary |
Decision | decision_id, owner, evidence_refs[], result, outcome_date | No evidence reference — records as an unevidenced decision |
The gate on AnswerIssued is the one that changed behaviour.
Making source version structurally mandatory meant a verbally-resolved answer could no longer enter the record
informally — and that, rather than any policy, is what moved resolution into the workbench.
Reducing the screening error rate
Q2 Reducing the screening error rateFive mechanisms, in the order they contribute. Only the third involves a model reading a patient note, and it's the narrowest possible use of one.
| Mechanism | What it does | False screen-fail | False screen-pass | Stale criterion |
|---|---|---|---|---|
| 1 · Structured criteria at the current version | The check runs against machine-readable criteria resolved for this site today, not a printed sheet | Reduces | Reduces | Eliminates |
| 2 · Deterministic rules for what is deterministic | Age, lab values, dates and windows are arithmetic. No model, no variance, no hallucination surface | Reduces | Reduces | — |
| 3 · Criterion-level reading of clinical notes | One targeted question per criterion against authorised notes. Finds evidence a human hasn't time to look for | Reduces most | Reduces | — |
| 4 · “Insufficient evidence” as a real outcome | Distinguishes “fails this criterion” from “I could not establish this”, and sends the second to a human | Reduces most | Reduces | — |
| 5 · Reason captured at the decision | Short controlled list, chosen by the person who made the call, in the moment | Makes visible | — | Makes visible |
Mechanism 5 reduces nothing. It's what turns an invisible error into a measurable one, which is why it was built first.
The evidence this approach works
Not speculative. In June 2024 — the same month as this engagement — Mass General Brigham published RECTIFIER in NEJM AI: a retrieval-augmented GPT-4 system answering thirteen eligibility criteria as separate questions against clinical notes, for a heart-failure trial. It's the closest published analogue to this architecture, and the result is instructive in a specific way.
Read the gap, not the headline.
Sensitivity barely moved — trained coordinators are already good at catching eligible patients. The entire gain is in specificity, more than ten points: the system was far better at correctly excluding people the staff wrongly flagged as potentially eligible. In practice that's screening effort spent on candidates who were never going to enrol. Per-criterion, published staff sensitivity ranged as low as two-thirds on the hardest criteria — which is the variance a structured, criterion-level approach removes.
Which lever actually moves the number
Move the levers
Machine-readable criteria, resolved to the version in force at this site today.
Criterion-level reading of notes, with insufficient evidence returned rather than guessed.
Doesn't reduce error. Determines whether you can see it at all.
eligible patient turned away
ineligible patient enrolled
per 100 screened
At today's settings nothing has changed — and the larger of the two error rates is one the organisation has no way of observing.
Where the human line sits, and why there
Every checkpoint is justified by something specific — an irreversible action, an asymmetric error cost, or an accountability a regulator requires. “Healthcare is sensitive” wasn't accepted as a reason for any of them.
Human decides
Clinical eligibility
Why: asymmetric error. A false pass exposes a patient to a trial they shouldn't be in, and clinical accountability can't be delegated to software.
Sees: criterion-level results, evidence found, evidence missing, version used.
Human decides
Ambiguous protocol interpretation
Why: resolving ambiguity changes what the protocol means for every site. That's authorship, not lookup.
Sees: both conflicting sources with effective dates, and how other sites were answered.
Human approves
Amendment impact map
Why: a missed site becomes a deviation, and reversal after sites are notified is expensive and public.
Sees: the full map with traversal coverage stated — including what the agent could not reach.
Human decides
Any contact with a site
Why: an unnecessary escalation damages a relationship the sponsor depends on for recruitment.
Sees: suggested priority, site history, comparable sites.
Human records
Screen-failure reason
Why: the value only means anything if it comes from the person who made the call, in the moment. A data-quality control rather than a safety one.
Sees: a short controlled list plus free text, inside the existing flow.
Human approves
Any model, prompt or corpus change
Why: a silent change alters clinical guidance at every site simultaneously.
Sees: the full evaluation run, including the version-correctness gate.
The four states the interface can be in
No numeric confidence score reaches the user. A named state tells someone what kind of checking to do; a percentage invites them to invent a private threshold. The escalation route is visible in all four states, including the confident one, so escalating never reads as the tool having failed.
Federated capability ladder — and where we drew the stop line why level 4 was deferred
| Level | Capability | What it answers | Prerequisite | Status |
|---|---|---|---|---|
| 1 | Local deterministic rules | Does this candidate meet each criterion? | Structured criterion object | Built |
| 2 | Local retrieval over authorised notes | Where is the evidence for this criterion? | Site-side deployment and local governance | Built |
| 3 | Federated analytics | Which criteria disqualify most candidates, across sites? Would this draft protocol recruit? | Site query capability, agreed query patterns, cohort thresholds | Piloted, two sites |
| 4 | Federated learning | Can we predict fit across site populations? | Stable features, validated labels, comparable semantics, evaluation design, privacy controls | Deferred |
| 5 | Multimodal federated learning | Does combining local modalities improve the decision? | Everything in level 4, plus multimodal data and validation at each site | Not pursued |
Deferred rather than declined. The published healthcare literature is consistent about why federated learning isn't a starting point: generalisation across heterogeneous sites, residual privacy leakage, communication cost, governance and evaluation design all remain live problems in real deployment. Levels 1 and 3 answer the questions this organisation actually has, at a fraction of the burden — and level 3 is what makes Track C possible.
Spotting errors and trouble on site
Q1 What's wrong on site, and where AI helpsOnce questions, answers and decisions exist as events, eight things become detectable that previously were not. None of them need a model. All of them need the event record from section 17.
Signal 01
Question clustering by criterion
One criterion generating far more questions than its peers, across multiple sites.
What it meansThe criterion is ambiguous or badly translated. A protocol defect, not site incompetence — and it is the single most actionable signal in the set.
Signal 02
Answers issued from a superseded version
Direct and countable, because every answer record carries the version it came from.
What it meansA propagation failure. Tells you which site, which criterion, and which day the drift started.
Signal 03
Override clustering
Where humans consistently disagree with the system, grouped by criterion and by site.
What it meansEither the system is wrong about that criterion, or the site is. Both are worth knowing and the decomposition tells you which.
Signal 04
Reopen rate
Questions answered, closed, then asked again by the same site.
What it meansThe answer did not land — unclear, not trusted, or not reaching the person who needed it.
Signal 05
Time to first correct answer after an amendment
Measured per site, from the effective date to the first answer citing the new version.
What it meansA direct measure of amendment propagation, per site, which nobody has today. This is the France hypothesis, made testable.
Signal 06
Criterion fail rates deviating from peers
A site failing candidates on one criterion far more often than comparable sites.
What it meansEither a real population difference or a misapplication. The country breakdown from Track C separates the two.
Signal 07
Question spikes preceding a deviation
A cluster of questions on one criterion, followed weeks later by a filed protocol deviation at that site.
What it meansA leading indicator. Once you have enough history, the spike becomes a prompt to intervene rather than a fact discovered afterwards.
Signal 08
Silence after an amendment
A site that asks nothing at all in the weeks after a substantial change.
What it meansUsually the worst signal in the set. A site asking many questions is engaging with the change. A site asking none has either not received it or is not reading it.
The absence of interaction is the signal nobody instruments.
Every operational dashboard in this industry counts activity. A site that goes quiet after an amendment produces no activity at all, so it produces no alert, and it is the site most likely to be running on the old criteria six weeks later.
What the country manager sees on a Monday
Rank sites by
Modelled Eight representative sites, to show how the ranking moves.
Proving it
How the exhaust becomes insight, what gets measured, what would stop the rollout, and which of our own assumptions did not survive.
What this is worth, and when it lands
Five layers of value, each traced from the problem through the mechanism to a measurement that would confirm it. Nothing here is a claimed saving. Each row names the number that would prove it and the baseline that has to exist first.
Bar length shows how much of the layer is available without further prerequisites, not its size.
A one-sentence answer takes three to nine days. Coordinators reconcile versions, chase owners and rebuild trackers. Most resolution happens by phone and is never recorded.
Version resolved before retrieval; answers carry their source; questions route to a named owner; the same answer publishes to every site that asked. Recording is a by-product rather than a task.
Time from question raised to answer used. Reopen rate. Escalation hops per issue. Share of answers carrying a current source version. Reconciliation hours removed per site per month.
Eligible candidates are turned away because evidence sat in an unread note or the criterion applied was superseded. That error is invisible: the sponsor has no lawful route to records for anyone who never enrols.
Criterion-level assessment inside the site, with insufficient evidence returned as a first-class result rather than a fail. Reason captured by the person who made the decision, at the moment they made it.
False screen-fail rate against a re-adjudicated baseline. Candidates recovered per 100 screened. Screen-failure reason latency. Time from identification to enrolment.
The protocol is written by people who cannot see the populations and executed by people who can, with no moment where those two facts meet. Reality arrives later as a change order.
A draft protocol run as a federated query against real site populations before lock. Amendment blast radius computed from the protocol graph and shared with the CRO before anything is priced.
Substantial amendments per study, against the sponsor's own historical rate. Pre-lock prediction against actual recruitment at close-out. Criteria changed after lock for recruitment reasons.
An answer from a superseded amendment looks correct in the audit trail. Nothing records which version informed a decision, so a version error is undetectable after the fact.
Source version structurally mandatory on every issued answer. Deterministic version gate ahead of retrieval. Every override records a reason. Cross-source disagreement raises an alarm instead of being ranked away.
Protocol-version findings at monitoring and audit. Deviations attributable to stale criteria. Share of decisions carrying an evidence reference. Detected extraction defects per amendment.
Each study starts from scratch. What one country learned about a criterion does not reach the next protocol, and the sponsor–CRO relationship resets at every kick-off.
Criterion-level question volume, screening outcomes and feasibility predictions accumulate across studies. Close-out compares prediction with actual, and the delta feeds the next protocol's design.
Feasibility calibration over time. Reuse of criterion structures across studies. Proportion of protocol decisions supported by prior evidence. Recurrence of the same issue across studies.
When each layer lands
The layers do not arrive together, and presenting them as though they do is how these programmes lose credibility in year two.
| Layer | Baseline required first | Owner of the number | Confounder to control for |
|---|---|---|---|
| Operational | Query log analysed; current resolution time measured rather than estimated | Clinical operations | Study phase — activity volumes differ hugely between start-up and steady state |
| Recruitment | Re-adjudication study on a sample of past screen failures | Clinical operations with medical | Indication and site population mix; seasonality in referral |
| Protocol | Historical amendment rate per study, with reasons classified | Clinical development | Regulatory-driven amendments, which no design change prevents |
| Quality and risk | Current findings rate from monitoring and audit | Quality assurance | Detection effort — more looking finds more, which looks like getting worse |
| Strategic | Nothing. It starts accumulating from the first study and cannot be backfilled | R&D leadership | Attribution across a portfolio that is changing for other reasons |
The cheapest thing we built was a version resolver. The most valuable thing it enabled was a conversation the sponsor and the CRO had never been able to have.
That is the shape of value in a regulated enterprise. It rarely arrives as automation of an existing task. It arrives when a decision that used to be made blind can be made on evidence, and the technology that makes that possible is usually smaller and duller than the decision it unblocks.
From recorded events to decisions someone can act on
Once questions flow through one surface, the organisation starts producing operational data it has never had. Here's how that becomes something a country manager can act on — one real question, traced all the way up.
Event
One thing that happened, recorded as a by-product of the work
Metric
Events counted against one owned definition
Signal
A metric that departs from what comparable units do
Explanation
The driver, decomposed — not a score, a reason
Decision and outcome
Someone acts, and the result links back to the action
Before — unanswerable, so not asked
- Which criterion is confusing the most sites?
- Is France slow because of capacity, or because amendments land late there?
- How often is a question answered from a superseded version?
- Which criteria disqualify the most candidates, and is that intended?
- Would this draft protocol actually recruit?
- Does the CRO's issue volume reflect site difficulty or CRO practice?
- Did the intervention we ran last quarter change anything?
After — answerable from data the work already produces
- Question volume by criterion, site, country and protocol version
- Time from amendment effective date to first correct answer at each site
- Share of answers carrying a current source version
- Criterion-level screening outcomes, aggregated above a cohort threshold
- Federated pass rates per criterion, per country, before lock
- Recurrence of the same issue across sites, from the CRO event feed
- Before-and-after on any metric, matched on study phase
What we measure, and what would make us stop
The previous pilot reported one number — 64.9% adoption — meaning a single authenticated login. Two levels down, use that had actually changed how work was done stood at 4.1%. Nothing measured whether the answers were right.
| Measure | What it is | How | Status |
|---|---|---|---|
| Version correctness | Every cited source is applicable to this site, country and date | Automated against the version object, on every release | Hard gate |
| Abstention quality | Abstains when it should — and doesn't when it shouldn't | Adversarial set: conflicts, gaps, superseded content, unanswerable questions | Hard gate, both ways |
| Groundedness | Every claim traces to a retrieved passage | Automated claim-to-passage binding, plus sampled human review | Threshold |
| Citation validity | The cited passage actually supports the claim | Automated, plus expert review on a sample | Threshold |
| Answer correctness | Expert-judged factual correctness | Stratified test set across studies, countries and languages | Threshold |
| Escalation correctness | Cases needing judgement actually reach a human | Human review of a sample of answers that were not escalated | Weighted by consequence |
| Feasibility calibration | Did the pre-lock prediction match actual recruitment? | Prediction-vs-actual per criterion, per country, at close-out | Monitored — Track C |
| Outcome effect | Screening error rate, resolution time, rework, amendment count | Cohort comparison against the re-adjudication baseline, matched on study phase | Decides scale |
Adoption, redefined so it can't flatter anyone
| Rung | Definition | Can't be faked by | Who acts on it |
|---|---|---|---|
| 1 · Provisioned | An account exists | — | Nobody. Not a metric. |
| 2 · Activated | Authenticated at least once | Provisioning | Nobody. The retired headline. |
| 3 · Repeat | Returned in a later week unprompted | A launch email | Product |
| 4 · Embedded | Used inside a task that produced an operational event | Browsing | Operations |
| 5 · Correct | The task completed correctly, judged against the harness | Frequent use of a wrong answer | AI governance |
| 6 · Outcome | Screening error rate or cycle time moved against baseline | Everything above it | Clinical operations leadership |
Rungs 5 and 6 didn't exist in the pilot. They're the only two that can support a decision to scale, and both depend on the re-adjudication baseline being established first.
Assumptions we tested before building
Section 04 tested the organisation's assumptions about the problem. This tests ours about the solution. Each row is something we believed at the point of designing, how we checked it, and what we did when it did not hold.
“Scaling the existing assistant to all 41 studies is the fastest route to value.”
“A central patient data lake would make screening questions answerable.”
“The deterministic criteria could be decided automatically.”
“An enterprise RAG platform the client already licenses could do this.”
“Federated learning is how you handle distributed clinical data.”
“A CRO scorecard would improve the relationship.”
“A dashboard would give leadership what it needs.”
“Adoption is a training problem.”
“Structuring the protocol is the foundational dependency.”
“The site-side pattern can carry more than screening.”
Still open
Six things this work did not settle. They are listed because a portfolio piece that resolves everything is describing something that did not happen.
Open What the query log actually contains
The pilot query log was never analysed. What people actually asked is the most informative dataset in the case and it remains unread. Every statement here about question types comes from interviews.
Open True volume of verbal resolution
Phone-resolved answers are invisible by definition. The claim that they are the majority comes from interviews, not measurement. The event model will size it, after the fact.
Open France's causal structure
Four candidate causes identified, none ruled out. The comparative analysis was designed; its result is not in this piece, and no intervention should be read as validated.
Open Site-side deployment at scale
Local computation assumes each site can host or reach a governed service. Site IT heterogeneity was never surveyed across the estate. Two pilot sites is not a portfolio answer, and Track C depends on this scaling further than anything else does.
Open Corpus ownership at scale
Structuring the protocol reduces curation effort without eliminating authorship. Who maintains criterion structure across 41 studies, inside which existing role, is decided in principle and not in practice.
Open Whether the CRO will engage with Track C
It assumes the CRO wants earlier, better-evidenced conversations. Some of its revenue comes from change orders. We designed the mechanism and did not test the incentive.
Ordered by dependency, not by quarter
No dates. Each step is gated by an artefact the previous step has to produce, which is what lets the sequence survive a slipped quarter, a reorganisation or a change of priority — none of which a phased plan survives.
Establish what's actually true
Re-adjudicate a sample of screen failures for a baseline error rate. Analyse the query log nobody had read. Reconcile the three screening counts to one owned definition. Decompose the France pattern instead of assuming capacity.
The protocol becomes data
Criterion extraction with schema validation, canonical study and site identifiers, and the version resolver. The smallest component in the architecture and the one everything depends on.
The workbench
Version-scoped answers with citation, deterministic routing to named owners, the four evidence states, and the evaluation harness running from day one rather than bolted on later.
Screening at two sites
Deterministic rules plus criterion-level note reading, inside the site environment, with the DPO-agreed output contract and cohort thresholds in place before the first patient is assessed.
The CRO event feed
Five event types on an agreed cadence, against the renegotiated data rights clause. Turns the quarterly review into a shared diagnostic.
Track C · the joint feasibility loop
Federated criterion queries on draft protocols, the joint working agreement, and the amendment blast-radius map shared with the CRO before anything is priced.
Signal detection and causal work
Anomaly models over site behaviour, targeted causal analysis on country differences, wider agent coverage.
Distributed learning, if it earns it
Federated learning only if a specific question requires it and the prerequisites in the ladder are met.
What people actually see
Three screens, placed last. The interface is the smallest part of this work and the part that would have been designed first if we'd got the sequence wrong.
Why the shape never varies
A user who sees the same eight parts every time learns where to look. The second time they see a conflict state, they already know part five is where the system tells them how much to trust it.
Why the route is always shown
If escalation only appears when the system fails, escalating becomes an admission that the tool did not work. Showing it every time makes asking a person an ordinary move rather than a defeat.
Why there is no confidence score
A percentage invites every user to invent a private threshold, which is the opposite of calibration. A named state tells them what kind of checking to do.
Why feedback sits in the answer
Corrections captured at the point of use are the only realistic source of new evaluation cases. Ask for them later and you get nothing.
Question
Can we include a patient whose HbA1c was taken 9 weeks ago?
No. The version in force at your site requires the value to be taken within 8 weeks of the screening visit. A repeat test is needed before eligibility can be assessed.
Question
Is prior GLP-1 exposure exclusionary if it stopped 14 months ago?
Two approved sources for this study give different washout periods and both are currently in force for your country. This is a protocol interpretation question, not a lookup.
Says 12 months · effective 04 Mar 2024
Says 18 months · effective 22 Mar 2024 — later
Candidate assessment · 13 criteria
Criterion-level results, not a verdict
evidence
The most capable component in the finished design is a two-hundred-line deterministic resolver that can say which protocol version applies to this site, today.
Everything else became possible once it existed, including the one thing neither the sponsor nor the CRO had ever been able to do: find out whether a protocol would work before agreeing to run it.