Blog / B–01 / Research direction
Why I Call It a Virtual Human
A model with a parameterized representation of one person’s measured biological state could eventually help compare possible responses before deciding which interventions warrant laboratory or clinical testing. Our RNA work addresses some of the representation and validation problems such a model would face.
A useful model of a person should estimate response to a defined biological change, state what it does not know, and survive tests on new measurements.
The phrase “virtual human” is easy to misread. It can sound like an avatar, a digital character, or a complete copy of a person. I mean something narrower and more demanding: a computational model conditioned on measurements from one person, built to answer a defined biological question. I call it a virtual human because it is anchored to an individual rather than only a population average, and because proposed changes can be explored computationally before they are tested in that person.
The model is defined by the question
A virtual human should represent selected aspects of a person’s biological state and estimate how that state may change after a perturbation. A perturbation is a defined intervention or exposure, such as a medicine, nutrient, supplement, change in sleep or activity, or environmental factor. Disease progression is a related trajectory-prediction problem.
The word “selected” matters. No available measurement captures a whole person. A model built from blood RNA may say something about immune activity at one time and under one collection protocol. It does not contain the person’s full physiology, medical history, environment, or future. The model has to show that boundary rather than hide it.
VHL does not operate a virtual-human system for individual prediction or clinical use today. The virtual human is the research target. The work now is to identify which representations, data, perturbation records, and validation methods would make that target scientifically defensible.
The test is not whether a model sounds like a person. The test is whether it can predict a measurable change in that person’s biological state, with uncertainty stated in advance.
From retrieving knowledge to conditioning an answer
For most of my education, scientific authority arrived in fixed forms: textbooks, papers, protocols, and guidelines. Search changed how we reached that material. The evidence stayed where it was, but retrieval became immediate.
Learned models changed the interface again. A person can now ask a conditional question and receive a synthesized answer rather than a list of documents. The answer may be wrong. It may be confident for the wrong reason. I nevertheless see people using these outputs as working answers rather than treating them only as search results. It has changed what I expect a knowledge interface to do: answer a conditional question, not only retrieve sources.
I expect biology to move toward a related interface. A biological model will not escape population data or prior evidence. It will learn relationships from those sources, then condition an estimate on measurements from an individual. Instead of asking only, “What usually happens after this intervention?”, we may also ask, “Given this observed state, which responses remain plausible?”
The analogy stops at the interface. Language-model use does not establish that a biological prediction is reliable. A biological estimate must be falsifiable, calibrated for its intended use, and tested against observations collected after the prediction.
Biology becomes useful under change
A static measurement can identify differences from a defined reference, pathways that differ between groups, or cell states present in a sample. Those are useful observations. They do not answer the question that motivates me: what happens next if something changes?
Estimating response requires temporal evidence and, when the claim concerns an intervention, a design that supports causal inference. Two people with similar measurements can respond differently because of prior exposures, genetics, tissue context, medication, behavior, or variables we did not collect. An individual-conditioned model must distinguish a plausible association from a counterfactual claim about an intervention.
The output should not be a single definitive future. It should be a set of possible responses, their time horizons, the uncertainty around them, and the conditions that would make the estimate unreliable. A model that cannot detect missing context or distribution shift should not be used as if it can.
Questions I want to make testable
I am interested in applications because they force the model to answer a concrete question. They are research targets, not services VHL offers now.
Which response differences should a prospective study test?
A modeled cohort could explore heterogeneity, identify candidate subgroups, and help prioritize trial hypotheses. It cannot establish safety or efficacy and cannot replace human trials.
Which treatment hypotheses are plausible enough to test?
An individual-conditioned estimate could narrow hypotheses for further laboratory or clinical evaluation. It would not by itself select a medicine or determine care.
How does measured state change after nutrition or behavior changes?
Repeated measurements collected around a defined exposure could characterize change after diet, supplements, sleep, or exercise. Causal claims would require an appropriate experimental or quasi-experimental design.
Which paths of progression, recovery, or relapse remain possible?
Repeated measurements may constrain possible trajectories and intervention windows. They will not make a person’s future deterministic.
Why our work begins with RNA
RNA abundance provides a time- and tissue-dependent measurement of gene expression. In bulk samples, it also reflects cell composition. That makes it a useful place to study biological state, but it is noisy and incomplete.
PBISC asks how much cell-level structure a bulk PBMC RNA profile can support. Bulk datasets exist for many cohorts in which single-cell measurements were not collected. If a model generates pseudo-single-cell structure from bulk measurements, every downstream claim still has to be tested separately. Generated cells are not measured cells.
RNA-LLM asks a different question: can a transcriptomic profile become a learned input to a language model without serializing thousands of gene values into text? The project studies how RNA state can condition language-model outputs and downstream tasks. It does not turn an RNA profile into a complete model of a person.
Our leakage-controlled evaluation work asks whether performance transfers to new donors and studies. This matters because a model can appear to understand biology while recognizing cohort, batch, site, or repeated-subject signals. A virtual-human research program cannot be built on that shortcut.
Neither PBISC nor RNA-LLM is the endpoint. They address resolution and representation problems that we can work on now. A credible individual model will likely need longitudinal molecular measurements, physiology, clinical context, behavior, environment, intervention records, and observed outcomes. We do not yet know which combination will be sufficient for which question.
A biological model has to earn trust
It is tempting to describe biological model error as hallucination. More precise failure modes already exist: confounding, poor calibration, missing variables, data leakage, distribution shift, and unsupported extrapolation. Each one calls for a different test.
I would not trust a model because its output is fluent or biologically plausible. I would look for external cohorts, prospective predictions, pre-specified endpoints, calibration, comparisons with simpler baselines, and failure cases reported alongside successful results. I would also ask whether the model knows when the input falls outside the data it learned from.
Clinicians would have more work to do than detect a wrong output. Any future clinical use would require them to judge applicability, supporting evidence, interactions, safety, comorbidities, and patient preferences. A model may organize evidence or estimate a response. It does not carry clinical responsibility.
I am not asking anyone to trust a virtual human today. The research question is what evidence would justify taking its output seriously for a defined task. VHL begins with RNA because it gives us a tractable set of measurement, representation, and evaluation problems. The program extends beyond RNA, but every step requires evaluation on new donors, cohorts, interventions, and measurement types.