Blog / B–02 / Clinical development
The Age of Trusting Parameters
Where virtual humans could enter drug development
Clinical trials estimate effects across groups, while treatment response depends on biological context. Models conditioned on individual measurements could help explore that heterogeneity—if their predictions earn a defined role through prospective evidence.
The pressure to model individual response is real. Adoption should depend on context of use, calibration, prospective validation, and the cost of a wrong decision.
Modern clinical trials do something indispensable: they estimate whether an intervention has an effect in a defined population. Their central result is usually a group-level estimate. That result can be rigorous and still leave a second question unresolved: what will happen in a person whose biology differs from the population represented by that estimate?
The limits of the average patient
Age, genotype, organ function, immune state, microbiome, comorbidities, prior exposure, and concomitant medicines can all alter treatment response. Trials account for some of this through eligibility criteria, stratification, enrichment, covariate adjustment, subgroup analysis, and adaptive design. Those methods matter. They do not make the number of clinically relevant combinations small.
The phrase “average patient” is therefore shorthand for a statistical problem, not a criticism of trial participants or randomized trials. Most confirmatory studies are powered around a population-level estimand. Important heterogeneity can remain beneath that estimate, especially when each subgroup is small or when several variables interact.
When a medicine produces less benefit or more harm in practice, the cause is not always a defect in the molecule. The treated person may differ from the development population in a way the study could not resolve. Historically, many such differences were treated as residual variability. Better measurement turns at least some of that variability into variables that can be tested.
The problem is not that clinical trials use averages. It is that no feasible trial can prospectively test every biological context in which a treatment will later be used.
A model system can be useful and still fail to transfer
Medicine has always relied on biological proxies. Pasteur developed rabies post-exposure inoculation through work in dogs and rabbits before his team treated Joseph Meister in 1885. The Toronto insulin work moved from depancreatized dogs to pancreatic extracts for patients in 1922, with purification by James Collip proving decisive. Mouse infection experiments in 1940 were followed by cautious human testing and therapeutic use of penicillin in 1941.
These were not careless substitutions of “animal” for “human.” They were early translations made with the evidence and methods available at the time. Their success established the value of model organisms. Later failures established the boundary: a useful proxy is not a replica.
Thalidomide made that boundary visible. The drug was marketed in West Germany in 1957 and withdrawn in 1961 after its association with severe congenital malformations became clear. Standard mice and rats are comparatively resistant to its teratogenic effects, whereas susceptible rabbits, non-human primates, and humans can show limb defects. The explanation is not one metabolic pathway alone; species differences in molecular targets, degradation pathways, pharmacokinetics, dose, and developmental timing all matter. The disaster added decisive public momentum to the 1962 Kefauver–Harris Amendments, which tightened investigational-drug control and required substantial evidence of efficacy. It did not, by itself, create the modern Phase I–III sequence.
TGN1412 exposed a narrower but equally important gap in 2006. Six healthy volunteers received 0.1 mg/kg of the CD28 superagonist, one five-hundredth of the no-observed-adverse-effect level used to set the starting dose from cynomolgus-monkey studies. All six developed a rapid systemic inflammatory response and critical illness. Later work found that CD4+ effector-memory T cells express CD28 in humans but not in the macaque species used for preclinical testing, helping explain why the primate data missed the human response. That difference is one part of a more complicated mechanism, not a universal indictment of animal studies.
Both cases support a limited conclusion: apparently reassuring evidence can fail when the model system omits a response-relevant feature of the target population.
A related generalization problem exists within our species
Human beings are closer to one another than different species are, but clinically consequential differences remain.
Codeine requires CYP2D6-mediated conversion to morphine. People with an ultra-rapid-metabolizer phenotype can produce dangerous morphine concentrations even at labeled doses. After reports of deaths and life-threatening respiratory depression in children following tonsillectomy or adenoidectomy, the FDA added a boxed warning and a contraindication for that postoperative setting in 2013. Genotype was important, but so were age, obstructive sleep apnea, dose, and clinical context.
Abacavir provides a more constructive example. Its multi-organ hypersensitivity syndrome is strongly associated with HLA-B*57:01. In PREDICT-1, prospective screening and exclusion of carriers reduced immunologically confirmed hypersensitivity from 2.7% to 0%; clinically diagnosed reactions fell from 7.8% to 3.4%. The biomarker did not fully determine response: its positive predictive value was 47.9%. It made one source of risk measurable enough to change practice.
These examples do not show that every adverse event can be predicted from a single marker. They show the opposite. Codeine and abacavir were tractable because a large part of the risk could be connected to a defined molecular feature. Polygenic effects, microbiome composition, interacting medicines, immune history, and time-varying physiology create a much larger state space.
Precision creates a granularity problem
The theoretical response is to keep stratifying: test by genotype, age, comorbidity, prior treatment, organ function, environment, and every interaction among them. In practice, the combinations grow faster than prospective cohorts can support.
Clinical development already uses enrichment strategies, adaptive and platform designs, covariate models, external controls in limited settings, and n-of-1 approaches. These methods reduce waste and answer questions that a single fixed design cannot. None makes exhaustive multidimensional stratification practical.
Cost is part of the constraint, but it should be stated carefully. One analysis of 138 pivotal trials supporting FDA approvals in 2015–2016 estimated a median direct cost of $19 million, with a range from under $5 million to $346.8 million. Trial size, duration, endpoints, indication, and site structure change the number dramatically. Repeating even the median-sized study across many interacting strata would still be economically and logistically prohibitive.
Explore more of the heterogeneity in models
This is where virtual biology can enter drug development. A model conditioned on molecular, physiological, clinical, and exposure data could generate hypotheses or task-specific response estimates for virtual cohorts. It could explore heterogeneity before a protocol is fixed, rank hypotheses, identify candidate enrichment variables, compare dose or sampling strategies, and expose scenarios that deserve laboratory or prospective clinical testing.
More inputs do not solve sparse evidence by themselves. A credible model must learn structure that transfers—such as dose–response, temporal, cell-state, or mechanistic relationships—and then show on external or prospective data that its conditional predictions remain calibrated. If performance collapses on a new cohort, genotype, intervention, or measurement platform, a larger feature set has not solved the problem.
This is no longer outside the regulatory vocabulary. The FDA describes model-informed drug development as the use of computational modeling and simulation to integrate nonclinical data, clinical data, prior information, and biological knowledge. Its applications already include dose selection, trial design, therapeutic individualization, prediction of clinical outcomes, and assessment of safety mechanisms. The 2026 ICH M15 guidance provides a framework organized around the question of interest, context of use, model influence, consequence of a wrong decision, and model risk.
That framework is more useful than asking whether simulation is “real evidence.” A low-risk dosing question and a high-risk decision to expose a patient to a novel treatment cannot demand the same validation. For a defined question, model-informed analysis may support dose selection or reduce the need for an additional dedicated study. It does not establish the safety or efficacy of a new treatment by itself; human evidence remains central.
Where might response differ?
Map plausible heterogeneity and prioritize variables before committing participants and sites.
Which study would be most informative?
Compare eligibility, dose, timing, sampling, endpoint, and subgroup strategies computationally.
Does the model add prospective value?
Pre-specify predictions, collect outcomes, and compare performance with simpler baselines and existing practice.
How much weight should the output carry?
Match model influence and required evidence to the consequence of being wrong.
What LLM adoption actually tells us
Large language models changed the interface to information. A user can ask a conditional question and receive a synthesized answer from a parameterized model rather than opening a list of sources. That change matters because it shows that model output can become a routine starting point for inquiry.
By “trusting parameters,” I do not mean trusting model weights in the abstract or accepting every output as fact. I mean allowing the output of a parameterized model to carry some weight in a defined decision.
It does not show that people universally trust LLMs, or that such trust is justified. Experiments have found that users can overestimate LLM accuracy and struggle to distinguish correct from incorrect answers based on fluent explanations. Other experiments found no greater trust in LLM recommendations than in AI-selected search snippets. Use is widespread; trust remains contextual and often poorly calibrated.
The biological analogy therefore stops at interface and adoption pressure. A model may rank hypotheses in exploratory research; a regulated development decision needs validation tied to its context of use; patient care imposes still higher evidentiary and professional obligations. Across all three, fluent output is not enough. What matters is prospective performance, calibrated uncertainty, explicit failure conditions, auditable inputs, and evidence that the model improves a defined decision.
The threshold is not accuracy alone
A model can score well and still be unsuitable for a clinical-development decision. The threshold depends on what the model is asked to do, how much influence its output receives, what evidence accompanies it, and what happens when it is wrong.
I expect economic pressure to push validated individual-response models into more parts of drug development. Direct empirical coverage of every relevant biological combination is infeasible, while biological measurements and computation continue to expand. That creates a strong incentive. It does not make adoption inevitable, and it does not make every virtual cohort credible.
The real question is narrower: can a model demonstrate prospective, calibrated value beyond existing trial methods for a defined task? If it can, the argument for using it will not rest on the phrase “virtual human.” It will rest on evidence that it improves a named decision while making better use of scarce human experiments.