Publication-partitioned splits and analysis code for the ImmunoStruct immunogenicity benchmark
收藏资源简介:
Evaluation splits and analysis code for the IEDB immunogenicity dataset released with ImmunoStruct (Givechian et al., Nature Machine Intelligence 8, 70-83, 2026, doi:10.1038/s42256-025-01163-y). The evaluation released with that model partitions peptide-HLA rows uniformly at random, so rows contributed by the same source publication can fall on both sides of a split. This deposit provides splits that instead keep each source publication intact, over exactly the 24,538 rows the published pipeline consumes. Contents data/splits/immunostruct_iedb_pubsplit.csv: peptide, allele and label as released, the recovered source publication identifier, a publication-grouped five-fold assignment in which no publication appears in two folds, and a leave-one-publication-out column. Folds are balanced on positive count rather than row count. data/derived/: the per-publication tables and exact null bands. analysis/ and figures/: seeded, deterministic code that rebuilds the splits and regenerates every figure. No GPU is required. patches/: two unified diffs against the ImmunoStruct repository that add a source-grouped split option. They change no model, no hyperparameter and no loss function. rerun/loso.sbatch: the job script for the leave-one-publication-out re-runs. Source attribution was recovered from the Immune Epitope Database (https://www.iedb.org) for every row. The peptide, allele and label columns correspond row for row to the file released with ImmunoStruct. Licensing differs by directory and the NOTICE file is authoritative. Data are CC BY 4.0, analysis code is MIT, and the two patches are derivative works of ImmunoStruct and remain under its Yale Non-Commercial License. Supporting material for a reusability report submitted to Nature Machine Intelligence.



