遇见数据集

AgenticGenomics: pharmacogenomics benchmark dataset for 'Validated skills are necessary but not sufficient for trustworthy agentic genomics'

收藏
Zenodo2026-08-16 更新2026-08-20 收录
官方服务:

资源简介:

Benchmark dataset for Corpas et al. (2026), "Validated skills are necessary but not sufficient for trustworthy agentic genomics" (under revision at Cell Genomics; CELL-GENOMICS-D-26-00551). Design. The package contains the matched five-configuration comparison (13,200 attempted evaluations across eight models, 110 CPIC Level A cases and three replicates; 13,199 records returned), the rule-extraction and bidirectional corrupted-contract experiments, GeT-RM external consensus comparisons, four-cohort real-genome analyses, and the input-normalisation experiment. Earlier three-arm datasets are retained and labelled as such. Version 1.7.0 (16 August 2026). This version supersedes 1.6.1 and exists to make the deposited code match the figures in the submitted manuscript. Five figure generators had drifted from the previously deposited source archive, so figures in the submission were not reproducible from the tag the manuscript cited. The source archive is replaced with agentic-pgx-benchmark-v2.4 (commit 2dce53a) and CHECKSUMS.sha256 is regenerated. No evaluation data changed: all 56 data files are byte-identical to 1.6.1, which is verifiable from the checksums in both records. What changed in the code. The material change is 82-figure7-real-genome.py. The curated-benchmark bar in Figure 3 was three literals, 0.96 / 0.0 / 0.04. Nothing in the data is 0.96 and the two round numbers beside it were placeholders, so the one bar every cohort bar is compared against was unreproducible from its own generator. It is now computed from the scored rows, giving 96.44 / 0.00 / 3.56, which is the 96.4% the manuscript reports. The other four generators carry a vocabulary change into the artwork, from "cell" to "configuration" for an arm of the matched comparison, because a check that reads manuscript text cannot see a word that lives in pixels. 81-validate-submission-package.py was rewritten: it validated a package superseded weeks earlier and held its tag and version DOI as constants, so it failed to run rather than failing loudly. Vocabulary. An arm of the matched comparison is a "configuration" throughout the manuscript and this record. Earlier versions of this description called it a "cell", which readers confused with a table cell and with the 336 lethal-class evaluations. The data files retain the column name cell and the script code/69-cell-provenance.py keeps its filename, because renaming deposited files would break the file listing of a published version DOI. The deposit therefore predates the vocabulary change, deliberately. Provenance. Each evaluation row records the model, the prompt inputs or the hash that freezes them, the raw response text, input and output tokens and cost in USD. No row carries a wall-clock timestamp or a provider response ID; these were not captured at run time and are not reconstructable. PROVENANCE.md states this explicitly rather than reconstructing an approximate time from file modification dates. What bounds the runs instead is an externally attested publication date recorded by Zenodo, combined with the provider billing period. Reproducibility. Code is pinned to commit 2dce53a and immutable tag agentic-pgx-benchmark-v2.4. CHECKSUMS.sha256 covers the data, release notes, provenance record and tagged source archive. Every quantitative claim in the manuscript is recomputed from this deposit by a registered-number check in the repository. Cohort data-access and redistributability limitations are documented in the repository; no new sequencing was performed.

提供机构:
Zenodo
创建时间:
2026-08-16
二维码
社区交流群
二维码
科研交流群
商业服务