Measured biology
Molecular, cellular, clinical, behavioral, and environmental observations collected with time and context.
Research / 2026
VHL’s long-term research program is to condition computational models on measurements from an individual and estimate how that biological state may change under a defined perturbation. Any such estimate would need quantified uncertainty and validation against observed outcomes.
Current work begins with RNA: recovering cell-state structure, representing transcriptomes, and building leakage-controlled benchmarks. These projects study components relevant to the long-term aim. They do not yet constitute a virtual-human model or a clinical decision tool.
Research model
A virtual human is not a complete digital replica. It is a task-specific, updateable computational representation of selected aspects of human biology, conditioned on measurements from an individual and evaluated for a defined question.
Molecular, cellular, clinical, behavioral, and environmental observations collected with time and context.
A computational representation that separates what was measured from what the model inferred or compressed.
A medicine, nutrient, behavior, exposure, or disease process specified with dose, timing, and assumptions where relevant.
A conditional estimate compared with new measurements, real outcomes, and explicit failure criteria.
Long-term research questions
These are research directions, not capabilities or services VHL offers today. Each would require longitudinal data, causal evidence, external validation, and prospective testing.
Compare modeled responses to candidate medicines, nutrition, supplements, sleep, exercise, or other defined exposures without turning an estimate into a personal recommendation.
Explore response subgroups and prioritize trial hypotheses before prospective human testing. Simulation cannot establish safety or efficacy or replace clinical trials.
Use repeated measurements to estimate possible progression, recovery, or relapse paths while exposing assumptions, time horizons, and uncertainty.
Current work / RNA
RNA is a dynamic but incomplete view of cellular activity. PBISC, RNA-LLM, and leakage-controlled benchmarks address resolution, representation, and testing questions relevant to broader response models.
PBMC bulk → pseudo-single-cell
PBISC asks how much cell-level structure an RNA profile can support. It generates an analysis-ready pseudo-single-cell count matrix from PBMC bulk RNA-seq. It is designed to work without a matched target-cohort single-cell reference at inference time.
Can large legacy bulk cohorts support cell-type and cell-state analyses that normally require single-cell measurements?
A generated count matrix for downstream immune-state analysis—not a claim that measured single cells have been recovered.
Cell composition, expression structure, pathways, and downstream biological conclusions must be tested separately.
Transcriptome → learned token
RNA-LLM tests an early bridge from a measured biological profile to a language-model interface. It compresses a single-cell, pseudobulk, or bulk transcriptome into one learned RNA token that conditions a language model on the measured transcriptomic profile.
Can a compact RNA representation preserve enough signal for biological reasoning across assay resolutions?
A learned bridge maps transcriptome features into the language model rather than turning thousands of gene values into a text prompt.
Biological plausibility, cross-study transfer, and evidence use matter alongside language quality.
Study-held-out / donor-held-out
A virtual-human model would need to transfer to new people and settings rather than recognize a familiar dataset. VHL develops evaluation protocols for autoimmune, infectious-disease, and vaccine-response transcriptomes that separate donors and studies before performance is measured.
Samples from the same cohort can share collection, processing, site, and batch signals. Random sample splits may reward those shortcuts.
Repeated measurements from one person must stay in one split so the test does not reward subject recognition.
Results should state the split unit, cohort provenance, endpoint, uncertainty, subgroup behavior, and known sources of leakage.
Evidence threshold
A convincing output is not enough. The full chain from observation to response must show provenance, uncertainty, transfer, and the conditions under which the model should not be used.
Record assay, context, cohort, donor, time, and access conditions.
State what is measured, inferred, compressed, generated, or missing.
Define the intervention, assumptions, time horizon, and uncertainty.
Test transfer, calibration, failure modes, and prospective performance.
Research inquiries
Include your affiliation, research context, available evidence, and the response or perturbation you want to study. Do not send patient information, controlled-access data, or confidential material.
Contact VHL