← Verity

Trust & methodology

Verity's whole premise is that you should not have to trust it. Every output is the return value of a cited, versioned engine you can recompute yourself. Here is exactly how, and exactly where the boundaries are.

The engines (and their sources)

Variant classification.

  • ACMG/AMP classification: the Richards et al. 2015 (Genet Med 17:405) combining rules, transcribed rule-by-rule, plus the ClinGen-endorsed 2018 SVI correction to the Likely-pathogenic rule.
  • Points model: the Tavtigian 2018/2020 Bayesian point system (doubling scale; posterior probability under a 0.10 prior). Run as a second, independent verdict; disagreement is surfaced, not resolved.
  • PVS1 strength: the Abou Tayoun 2018 loss-of-function decision tree (returns “unmet” rather than guessing when inputs are missing).
  • In-silico calibration: Pejaver 2022 calibrated REVEL thresholds for PP3/BP4 (an indeterminate score yields no criterion).
  • Population frequency: the Whiffin 2017 maximum-credible-allele-frequency framework for BA1/BS1/PM2.
  • Substitution severity: the Grantham 1974 amino-acid distance (from the composition / polarity / volume table), returning the canonical published matrix; surfaced for any missense.

Pharmacogenomics.

  • Star-allele phenotyping: CPIC/PharmVar diplotype → activity score → metabolizer phenotype, with the standardized CYP2D6 boundaries (Caudle 2020) and the CPIC diplotype-phenotype tables (Lee, Birdwell, Relling, Cooper-DeHoff, Gammal). Copy number that can't be inferred from SNVs yields indeterminate, never a defaulted *1/*1.
  • Warfarin dosing: the IWPC pharmacogenetic and clinical algorithms (Klein 2009, NEJM), term-by-term.
  • Recommendations: CPIC guidance is returned verbatim from a curated, per-row-cited table; the model never generates it. HLA risk is carrier status (a single copy contraindicates), not a metabolizer phenotype.

Protein.

  • ProtParam-class physicochemistry: isoelectric point (EMBOSS / Bjellqvist pKa sets, Henderson-Hasselbalch), molar extinction at 280 nm reported for both the oxidized and reduced species (Pace 1995), GRAVY (Kyte-Doolittle 1982), aliphatic index (Ikai 1980), and instability index (Guruprasad 1990 DIWV). Verified against ExPASy ProtParam known-answers.
  • Developability & synthesis: a sequence-liability scan and codon-optimized construct design (CAI, sliding-window GC, restriction-site repair, primer / Golden-Gate / Gibson design) built on the audited @/lib/cloning and @/lib/bioinformatics kernels.

Every engine is covered by known-answer unit tests against the published worked examples, and each was independently re-verified against the primary literature by an adversarial review, which has more than once caught a real transcription or off-by-one error before it shipped.

The case orchestrator

The Case Workspace plans a precision-medicine workup deterministically from the case inputs, runs each engine above, and (where raw evidence is supplied) derives the ACMG criteria itself and cites the derivation (a gnomAD allele frequency yields PM2/BS1/BA1; a REVEL score yields PP3/BP4). It surfaces prioritized safety flags (an HLA contraindication, a DPYD poor-metabolizer, a blocking validation error) and an evidence-coverage advisor that names the evidence types (functional, segregation, de novo) whose addition would resolve a VUS. The narrative brief quotes these already-computed, already-cited facts and introduces no number of its own; the resolved criteria that produced each verdict are shown in full.

What the model does, and never does

A language model may orchestrate a case, retrieve and cite evidence, and write plain-English explanations. It never assigns an ACMG code, produces a classification, or emits a number. Those come only from the deterministic engines above. An explanation is always rendered subordinate to the computed verdict and is stamped “explanatory only.”

Attribution, sealing, and falsifiability

  • Attribution. Every applied criterion must carry an evidence source. This is enforced in the engine and again in the database RPC: an un-sourced code is rejected, not silently accepted.
  • Sealing. Each saved classification is sealed with a SHA-256 hash over its evidence and verdict, on a hash-chained, PHI-aware audit ledger that records what changed and when, never the genomic values themselves.
  • Falsifiability. Anyone can re-derive a record from its stored inputs. When a verdict does not reproduce, Verity refutes it. Discordance between the two engines is preserved as recorded dissent.
  • Review. A classification is locked and requires molecular-pathologist sign-off before release.

Data handling (read this)

Genomic content (gene and variant notation) necessarily reaches the model provider when the platform writes a narrative. That is disclosure to a processor, not de-identification, and we do not claim otherwise. Real-patient use is gated at runtime and permitted only under a signed Anthropic zero-retention / no-training BAA plus Supabase and Vercel BAAs; the default configuration is synthetic / consented-research only. Cases are pseudonymous by schema; identifiers (MRN, SSN, DOB patterns) are actively rejected at ingest.

Engine integrity: fingerprints

Beyond citing each engine, Verity fingerprints it: a SHA-256 over the engine's output across a fixed input grid, recorded in a manifest. A build-time lockfile fails if any constant or formula changes without a deliberate re-record, so a wrong number cannot ship silently: it is the engine-level complement to the per-record seal. The fingerprints below are recomputed on this page and checked against the recorded manifest.

ACMG/AMP variant classification✓ matches manifest
Categorical (Richards Table 5) + Tavtigian points from attributed criteria
acmg-2015-richards+tavtigian-2020@2 · sha256:ce527ffc577f8b5fbb083886…
Grantham substitution distance✓ matches manifest
Physicochemical distance between amino-acid substitutions
grantham-1974@1 · sha256:58c7261f8a85708c892f2b91…
Protein instability index (DIWV)✓ matches manifest
Guruprasad instability index from the dipeptide instability-weight table
protparam-instability@1 · sha256:3d40be70d31cb6c94444632b…
Warfarin IWPC pharmacogenetic dose✓ matches manifest
IWPC square-root weekly dose from demographics + CYP2C9/VKORC1 genotype
iwpc-2009@1 · sha256:26ac8b51b575fb41b6191241…
REVEL in-silico calibration✓ matches manifest
Calibrated REVEL thresholds → PP3/BP4 strength bands
revel-pejaver-2022@1 · sha256:1041d586869f6eb9465bc606…
Population-frequency framework✓ matches manifest
Whiffin maximum-credible-AF + observed-AF → BA1/BS1/PM2 mapping
whiffin-2017@1 · sha256:11e964f17f4b0897f6d0454b…
Pharmacogenomic phenotyping + CPIC guidance✓ matches manifest
Star-allele → activity score → metabolizer phenotype, and CPIC recommendation table
cpic-pharmvar@1 · sha256:ddf9bb409dd2e19fc2ece428…
Protein physicochemistry (ProtParam-class)✓ matches manifest
Isoelectric point, molar extinction (280 nm), GRAVY, aliphatic index, molecular weight
protparam@1 · sha256:82a68cb2c124e91f27168baa…
Molecular descriptors + aromaticity✓ matches manifest
Wildman–Crippen logP, Ertl TPSA, Lipinski/Veber drug-likeness, and Hückel aromaticity re-perception from a SMILES
molecule-descriptors@1 · sha256:7e5a1a1386c09610c7f69be8…
Literature evidence-strength (OCEBM × PubMed)✓ matches manifest
Deterministic 0–100 target evidence-strength from best study design, corroboration, recency, and contradiction — the authoritative score behind a Verity evidence report (the AI never produces it)
ocebm-2011-pubtype@1 · sha256:2e339c2bff6d0682d52e9341…
Variant clinical-evidence confidence (ClinVar × literature)✓ matches manifest
Deterministic consensus significance + 0–100 evidence-confidence for a gene/variant from ClinVar gold-star review status, submitter concordance, and quote-grounded literature — evidence-gathering, NOT ACMG classification
clinvar-stars-lit@1 · sha256:e54f8bf71aff3b21f67a23f2…
Therapeutic evidence-landscape (ClinicalTrials.gov × literature)✓ matches manifest
Deterministic 0–100 evidence-landscape strength for a drug × indication from clinical trial phase/breadth/results and quote-grounded literature — describes HOW MUCH evidence exists that the drug was studied, NOT efficacy or a clinical recommendation (literature direction is reported as context, never folded into the score)
ctgov-phase-landscape@1 · sha256:0f1d86fefcec6c5a1937972c…
Whole-body PBPK simulator✓ matches manifest
Deterministic perfusion-limited physiologically-based pharmacokinetic simulation of a compound through the real circulatory topology (venous → lung → arterial → organs, gut/spleen draining portally through the liver), giving per-organ and plasma concentration-time curves plus non-compartmental PK. Vascular states are blood-referenced and converted to plasma for reporting through an explicit blood:plasma ratio (default 1). Tissue partitioning uses the COMPLETE Poulin & Theil tissue-composition method (phospholipid terms, the fu_p/fu_t binding correction, and the separate vegetable-oil equation for adipose) on the human composition table; it is systematically low for moderate-to-strong bases, which the result flags say. Hepatic elimination can be parameterised three ways, most specific first: saturable Michaelis-Menten on unbound drug (Vmax/Km), unbound-driven linear intrinsic clearance (well-stirred), or a whole-organ clearance. Compound ADME inputs are supplied and are NOT derived here; physiology, the ODE solution and every reported metric are computed. Research use — it predicts exposure under the stated model, it does not establish a dose.
pbpk-perfusion-limited@8 · sha256:e658a6141dc8f5b0a7ffa547…
PBPK covariate individualisation + virtual population✓ matches manifest
Scales the whole-body simulation to a subject’s covariates (metaboliser activity through the fraction metabolised fm, Child-Pugh hepatic grade, renal function, body weight) and generates a SEEDED virtual population whose 5th/50th/95th-percentile exposure bands are reproducible from the seed alone. Also computes therapeutic-window residence. Covariates scale clearance only; between-subject parameters are currently sampled INDEPENDENTLY, so the bands do not represent covariate correlation.
pbpk-precision@2 · sha256:18fbad34fad21df84e542ec4…
Drug–drug interaction + genotype exposure bridge✓ matches manifest
Simulates a perpetrator compound’s own pharmacokinetics, then re-simulates the victim under the resulting time-varying hepatic clearance, reporting AUC and Cmax ratios and an FDA-threshold classification. The same clearance-scaling contract carries a star-allele diplotype through the CPIC activity score to a whole-body exposure prediction. The interaction is applied through the fraction of the victim’s hepatic clearance the affected enzyme carries (fm), so the AUC ratio is bounded by 1/(1−fm) as inhibition becomes complete rather than rising without limit. fm defaults to 1 — the whole hepatic clearance responding — which is an upper bound unless a victim-specific fm is supplied.
pbpk-ddi@3 · sha256:ca31e6365c03595ab0d37c75…

Honest capability boundary

Verity is a suite of deterministic, guideline-cited calculators over attributed inputs, not an autonomous pipeline. It does not auto-annotate a bare variant (that needs VEP/gnomAD/ClinVar/REVEL upstream), call star-alleles from a raw VCF (it phenotypes a diplotype you supply), predict protein structure, dock a ligand, or output an individual clinical outcome. Its dosing and phenotype outputs are decision-support, not a prescription. It is research-use-only, non-diagnostic, physician-in-the-loop, designed to display the reviewable basis for every output and never to drive a time-critical decision, which is how it stays within the FDA non-device clinical-decision-support boundary. Determinism guarantees reproducibility, not correctness of the underlying guideline tables; those are versioned, dated, hashed, and maintained.

The ACMG/AMP classification engine is scoped to germline sequence variants (Richards 2015). It does not implement the AMP/ASCO/CAP 2017 four-tier system used for somatic (tumor) variants (Li et al., J Mol Diagn 2017). A somatic specimen is flagged in the workbench, and its ACMG output must be read as germline-framework evidence, never a somatic clinical tier.

Open the workbench →