PlatonicMHC Pan-Allele Dataset: 60,000 peptide–MHC class I complexes across 120 human HLA alleles
收藏资源简介:
A balanced dataset of 60,000 peptide–MHC class I (pMHC-I) complexes spanning 120 human HLA alleles (HLA-A, -B, -C, -E, and -G), constructed to evaluate cross-allele generalisation in structure-based pMHC binding prediction. Each entry pairs a 9-mer peptide with an HLA allele and a binary binding label, together with a three-dimensional structural model generated using PANDORA. The dataset is derived from Tadros et al. (2025) and restricted to human alleles and 9-mer peptides, with 250 binders and 250 non-binders sampled per allele to ensure class balance and equal representation across alleles. It is partitioned into an allele-disjoint train–validation–test split (approximately 80–10–10) using hierarchical clustering of MHC pseudosequences by pairwise PAM30 distance, following the protocol of Marzella et al. (2024), so that test alleles are absent from training and evolutionarily distant from those seen during optimisation. This dataset accompanies the MSc thesis PlatonicMHC: Fast and Robust pMHC Structure Prediction via Equivariance (University of Amsterdam).The dataset contents are located under projects/0/einf2380/pMHCI_data/db1/ within the archive.



