GLP-1 Virtual Cohort v1
收藏资源简介:
This upload contains the full technical documentation and a free evaluation sample (498 patients, 6 branches, 104 weeks each — 310,752 weekly records) — not the full dataset. The complete dataset — 20,000 patients, 12,480,000 weekly records, 104 weeks of weekly follow-up across 6 scenario branches — is available at: https://sentineldata.com.ua/dataset/glp-1-virtual-cohort Sample contents: sample_patients_v1.parquet498 rows sample_weekly_as_treated_v1.parquet51,792 rows sample_weekly_switch_drug_v1.parquet51,792 rows sample_weekly_dose_variation_v1.parquet51,792 rows sample_weekly_early_discontinuation_v1.parquet51,792 rows sample_weekly_treatment_holiday_v1.parquet51,792 rows sample_weekly_no_treatment_v1.parquet51,792 rows CSV mirrors are not included in this sample; Parquet is the canonical format. A short guide on opening Parquet files (HOW_TO_OPEN_PARQUET.txt) is included. The sample is a real subset of the full dataset — drawn by stratified (per-molecule) random sampling with seed 42, exactly 83 patients per molecule (6 × 83 = 498), with complete weekly records in all 6 scenario branches for every sampled patient. This is an extract, not a re-generation: every row was identity-checked against the full dataset. Unlike a naive first-N-by-ID slice, this sample is balanced across all 6 molecules, so it is representative of the full cohort's molecule mix. Fully synthetic. Contains no real patient data. Every generative parameter is traceable to a published, verified source. 1. Overview GLP-1 Virtual Cohort v1 is a fully synthetic, longitudinal virtual patient cohort of 20,000 patients followed weekly for 104 weeks (12,480,000 weekly records) across six GLP-1 / GIP / amylin / glucagon-receptor molecules, built for obesity/metabolic real-world-evidence modeling, weight-trajectory and adherence research, and counterfactual / causal-inference methods development. Weekly trajectories are driven by a two-compartment pharmacokinetic model with population parameters, calibrated so simulated population means reproduce published trial endpoints exactly. 2. Molecules & Doses MoleculeDoses modeledPatients Semaglutide2.4 mg~1/6 (≈3,333) Tirzepatide5 / 10 / 15 mg~1/6 (≈3,333) Orforglipron6 / 12 / 36 mg~1/6 (≈3,333) Retatrutide4 / 9 / 12 mg~1/6 (≈3,333) Eloralintide1 / 3 / 6 / 9 mg~1/6 (≈3,333) CagriSema2.4 / 2.4 mg~1/6 (≈3,333) Assigned dose is drawn from each molecule's trial-observed maintenance-dose distribution. Baseline population (20,000 patients): age, sex, race, height, weight, BMI, waist circumference, HbA1c, prediabetes/T2D status, comorbidity count, hypertension, blood pressure — all trial-matched by assigned molecule. 3. What Is Modeled (Weekly) Pharmacokinetics: two-compartment concentration model (primary moiety + second moiety for CagriSema's cagrilintide component), weekly mean concentration and exposure fraction of steady state, driving the pharmacodynamic effect. Efficacy: weight (kg, % change from baseline), BMI, fat mass, lean mass, fat fraction — weekly. Glycemic control: HbA1c trajectory (T2D and non-T2D patients). Adherence: latent per-week missed-dose propensity; weekly missed-dose flag. Adverse events: five weekly GI AE flags — nausea, vomiting, diarrhea, constipation, fatigue. Discontinuation: permanent-discontinuation flag with cause (AE-related vs. three non-AE categories). 4. Scenario Branches — Same Patients, Six Counterfactual Worlds All six branches simulate the same 20,000 patients (identical latent traits, shared random numbers) — branch-to-branch differences are purely causal, not sampling noise. Built for switch/discontinuation/dose-change counterfactual and causal-inference method development without needing a real trial arm. BranchDefinition as_treatedFactual: assigned drug and dose, trial-level adherence (STEP 1 anchor: 81.1% PDC ≥ 80%), drug-specific discontinuation switch_drugAs-treated through week 26, then switched to a uniformly random other molecule (own titration restarted) dose_variationAs-treated through week 26, then maintenance dose re-drawn from the same molecule's trial dose distribution early_discontinuationAs-treated through week 26, then all patients stop; weight regains toward baseline (STEP 1 extension kinetics) treatment_holidayAs-treated except weeks 26–37 off drug, resuming at week 38 with re-titration no_treatmentNever treated — placebo-arm trajectory of the assigned molecule's own trial 5. Files & Schema FileRowsGrain patients_v1.parquet20,000one row per patient (demographics, baselines, latent traits) weekly_as_treated_v1.parquet2,080,000one row per patient-week, factual branch weekly_switch_drug_v1.parquet2,080,000one row per patient-week, switch-molecule counterfactual weekly_dose_variation_v1.parquet2,080,000one row per patient-week, dose re-draw counterfactual weekly_early_discontinuation_v1.parquet2,080,000one row per patient-week, early-stop counterfactual weekly_treatment_holiday_v1.parquet2,080,000one row per patient-week, treatment-holiday counterfactual weekly_no_treatment_v1.parquet2,080,000one row per patient-week, never-treated counterfactual Plus: validation_report_v1.csv (113 checks), run_manifest_v1.json (seed, versions, all calibration factors), parameter_table_v1.csv (310 sourced parameters), literature_registry_v1.csv (32 sources), generator.py + build_parameter_table.py (reproducibility), and 14 validation figures (PNG + SVG). All content in English; format is Parquet. 6. Validation — 113 / 113 Checks Pass CategoryChecksResult Weight efficacy vs. trial endpoints (15 drug–dose groups)15PASS — exact, by in-simulation calibration GI AE incidence vs. trial targets (60 groups)60PASS — within 0.003 of target Discontinuation per drug6PASS HbA1c change, T2D (semaglutide, tirzepatide 5/10/15)4PASS Responder shares ≥5/10/15/20% (semaglutide vs. STEP 1)4PASS — 0.863 / 0.691 / 0.505 / 0.320 Weight regain after stop (STEP 1 extension: 32.4% retained @52wk)1PASS — 0.324 Prediabetes→normoglycemia reversion (STEP 1: 84.1%)1PASS — 0.841 Nulls / shapes / finiteness / cross-stage coherence22PASS The generator is deterministic: python generator.py --n 20000 --out <dir> with the pinned package versions (python 3.11.13, numpy 2.1.0, pandas 2.3.1, pyarrow 21.0.0, scipy 1.15.0) reproduces byte-identical parquet files at seed 42. Calibration disclosure: published trial endpoints are treatment-policy (ITT) estimands that mix adherent patients, non-adherent patients, and dropouts. This cohort reproduces those population means exactly by an in-simulation secant-solved plateau correction and a calibrated response-multiplier distribution — consequence: a fully adherent on-treatment patient loses somewhat more than the trial mean, i.e. the exposure–response slope is an effective parameter, not a physiological one. All calibration factors are recorded in run_manifest_v1.json. 7. Provenance 310 generative parameters: 221 VERIFIED (live literature search, DOI/PMID/URL attached), 70 ESTIMATE_SOURCED (anchored to a real source, flagged), 19 DESIGN (modeling conventions, labeled), 0 unsourced — backed by a 32-source literature registry: STEP 1/2/5, SURMOUNT-1, SURPASS-2, ATTAIN-1, ACHIEVE-1, TRIUMPH-1/2, REDEFINE 1/2, eloralintide phase 2, population-PK studies, FDA label, real-world adherence studies. Full audit trail (search queries, verification dates) ships in parameter_table_v1.csv. 8. Known Limitations Retatrutide mean drifts to ≈−27% by week 104 vs. trial −30.3% (kinetics calibrated at week 80, per the modeling convention used). Eloralintide T2D HbA1c is weight-linked drift only — no eloralintide T2D trial exists to anchor to directly. The GI-AE exact-count calibration polish does not re-derive disc_cause afterward (≤2% of weekly flags affected). Orforglipron bioavailability of 79.1% is used per the 2025 ADME study despite older estimates of 20–40% (documented conflict, resolved in favor of the newer source). TRIUMPH-1/2 anchors are Lilly topline press-release values; full phase 3 publications are still pending. ATTAIN-1 is anchored to the treatment-regimen (not efficacy) estimand. Synthetic data: not suitable for clinical decisions; associations absent from the source trials should not be treated as real. 9. Use Cases Suitable for: obesity/metabolic real-world-evidence and weight-trajectory modeling, GLP-1/GIP/amylin adherence and discontinuation research, counterfactual and causal-inference methods development (switch, dose-change, holiday, and stop scenarios on identical patients), health-economics and outcomes-research (HEOR) modeling, ML training/evaluation on richly annotated longitudinal weight-loss data, and education in pharmacokinetic/pharmacodynamic modeling of incretin therapies. NOT suitable for: clinical decision-making — this dataset is fully synthetic and contains no real patient data. 10. License & Distribution Format: Apache Parquet (+ CSV/JSON/MD documentation). Delivery: One-time purchase; download link provided after payment. A free, stratified sample (498 patients, 83 per molecule, complete weekly records in all 6 branches) is available for evaluation before purchase.



