遇见数据集

TheraOT: a traceable, evidence-tiered benchmark of clinical and therapeutic CRISPR off-target editing

收藏
Zenodo2026-07-20 更新2026-08-01 收录
官方服务:

资源简介:

Benchmarking CRISPR off-target prediction methods requires a ground-truth dataset that reflects clinical reality: sparse validated sites per programme, heterogeneous assays, strict provenance, and the principle that un-interrogated sites are not negative. Existing resources such as CRISPRoffT (Wang et al. 2025) and TrueOT (Kota et al. 2021, preprint) provide valuable computational benchmarks, but do not annotate evidence tiers, regulatory-document provenance, or the not-reported--negative constraint needed for clinical interpretation. We present TheraOT (v0.1.0, local staged version), a curated, evidence-tiered benchmark of 72 entries spanning 10 therapeutic CRISPR programmes (plus a computational companion). Entries are stratified across four evidence tiers: C1 (patient-derived cells at the end of product manufacturing, classified as near-C1; under the strictest reading these are C2; see Section sec:tierdefs), C2 (patient-derived edited product cells, ex vivo), T1 (primary human cells at therapeutic dose), and T2 (proxy cell lines). TheraOT captures 4,786 strict method-rescorable off-target sites: 4,716 WGS off-target rows from the CRISOT benchmark (Chen et al. 2023 ), 69 GUIDE-seq off-target rows in HEK293 cells (Chen et al. 2023 ), and one exact sequence/coordinate/quantification row in the main CSV (the exa-cel/Yen CPS1/rs114518452 C2 clinical product row). The three iGUIDE/NYCE-TCR positive rows are source-documented near-C1/C2 findings but are not strict method-rescorable because current coordinates/sequences are reconstructed or approximate and read counts are missing. The 33 Yen et al. (2025) C2 entries provide hg38 coordinates and per-participant editing frequencies but, except for the sequence-bearing CPS1/rs114518452 site, lack off-target candidate sequences and are therefore not independently rescorable. Three auditable clinical findings emerge: enumerate*[label=()] a corrected enumeration-recall gap: after filtering is_on_target == False, Cas-OFFinder at an illustrative 4 mismatch setting omits 97.964% (4,620/4,716) of CRISOT WGS off-target positives and 47.826% (33/69) of CRISOT GUIDE-seq off-target positives; an in-silico read-support proxy deflates the WGS interpretation because 79.4% of the omitted WGS positives have <5 supporting indel reads (median 1 read, median indel frequency 1.79%) and only 373/4,620 (8.1%) have 20 reads, while NTLA-2001 retains a separate SITE-seq-only gap in which Cas-OFFinder and GUIDE-seq each missed 57.1% (4/7) of validated sites at 27 EC_90 supratherapeutic dose (zero off-target sites were detected at therapeutic dose 3 EC_90); the clinically confirmed exa-cel off-target, CPS1/rs114518452 variant-created site in 3/91 patients (SCD2: 0.45%, SCD20: 1.00% edited-product indel frequencies), supporting the population-conditional safety thesis; and the FDA PMR#2 regulatory mandate requiring per-continental-group, carrier-specific variant-aware re-analysis for Casgevy. enumerate* The CRISOT enumeration gap is therefore real but not, by itself, evidence that standard tools are dangerously wrong; terminal PRJNA921906 raw replay is still required before any safety interpretation. TheraOT's contribution is a corrected, deduplicated benchmark resource that supplies the evidence-tiered, regulatory-traceable inputs a downstream population-conditional certification framework (PEG-Cert) would require for explicit coverage claims; TheraOT itself does not compute or execute such a certificate. Within-positive Spearman correlation on the 4,716 WGS-validated sites shows all five scored methods near zero (||<0.09). Most 95% CIs span zero; exceptions are PCSK9-WT CRISOT/MOFF/CRISPRnet, BCL11A-opti MIT, and pooled CRISOT (=+0.030, 95% CI [+0.002,+0.058], p=0.042). Only the small PCSK9-WT CRISOT and MOFF effects survive Bonferroni correction over 20 by-guide tests, and these single-guide signals are not interpreted as general ranking performance; no genuine AUROC/AUPRC is computable because MOESM9 contains only validated positives. TheraOT is released as a versioned parser-safe CSV (theraot_benchmark_v4_clean_audit_fields.csv) with a per-row public-rescorability audit schema (six fields: scorable, public_rescorable, sequence_available, coordinate_available, quant_available, gated_reason) plus a provenance manifest under CC BY 4.0 terms, archived at Zenodo under a reserved DOI and mirrored in a version-controlled repository, released publicly upon publication. A separate Supplementary Information section documents the Cancellieri et al. 2022 CRISPRme candidate-enumeration companion without altering the 72-entry benchmark (67 C1/C2/T1/T2-or-near-tier rows plus 5 NOT_TESTED computational candidates).

提供机构:
Zenodo
创建时间:
2026-07-20
二维码
社区交流群
二维码
科研交流群
商业服务