Recursive confirmation bias in MHC class-I epitope prediction: model checkpoints, datasets and predictions
收藏资源简介:
Model weights and data supporting the recursive-data-corruption experiment reported in Resolution of recursive data corruption to transform T-cell epitope discovery. Five independent dataset versions, each with an iteration ladder in which training labels are rewritten by the previous iteration's own predictions, so that bias compounds. Thirty trained peptide-MHC ranking models result. Measured against the corrupted labels they were trained to agree with, the models appear to improve with every round; measured against clean labels they do not. Contents checkpoints.zip (2.3 GB) — 30 model checkpoints datasets.zip (393 MB) — 85 clean and corrupted training and validation bundles predictions.zip (548 MB) — 209 prediction tables SHA256SUMS.txt — checksums for the three archives Analysis code, and instructions for unpacking these archives, are at https://github.com/deepflare/IEDB_RCB. The repository also commits a small extract of the labels and scores, so the reported metrics and figure can be reproduced without downloading anything here. Underlying observations derive from the Immune Epitope Database and from previously published mass-spectrometry studies, which remain available from their original sources.



