Dataset and Benchmark Results for Conformational B-Cell Epitope Prediction Tools
收藏资源简介:
GEN416 This dataset contains GEN416, a non-redundant collection of 416 antibody–antigen complexes derived from SAbDab using stringent structural quality filtering and sequence-based redundancy reduction (CD-HIT, 70% identity). It serves as a curated structural dataset for evaluating conformational B-cell epitope prediction methods, including analysis of model generalization and potential training-data bias. The dataset includes PDB-derived structural information, antibody–antigen chain mappings, and residue-level epitope annotations defined using a 4 Å heavy-atom distance cutoff. BENCH53 This dataset contains BENCH53, a strict benchmark subset of 53 antibody–antigen complexes derived from GEN416 after removing any overlap with training data of 13 evaluated epitope prediction tools using 99% sequence identity filtering. It provides a leakage-free benchmark for unbiased comparison of conformational B-cell epitope prediction methods. The dataset includes PDB metadata, antibody–antigen chain information, residue-level epitope labels defined by a 4 Å distance criterion, and per-residue predictions from multiple tools.



