Cognate and non-cognate antibody-antigen complexes: 1.8 million AI-generated structures
收藏资源简介:
Abstract Antibody-antigen binding prediction remains a central challenge for AI-driven therapeutic discovery, particularly in discriminating cognate interactions from structurally plausible but incorrect pairings. We present a controlled, AI-method- and antibody-format-agnostic evaluation framework that measures binding specificity under realistic conditions. Using 106 experimentally determined single-chain antibody (nanobody)-antigen complexes and 11,342 shuffled non-cognate pairings, we benchmarked publicly-available state-of-the-art structure prediction methods (AlphaFold3, Boltz-2, Chai-1). Although the methods tested often generated geometrically plausible complexes, internal confidence metrics (ipTM) frequently failed to discriminate correct from incorrect pairings. Increased sampling improved structural refinement but not pairing discrimination, indicating that computational resources are better allocated across independent seeds and explicit negative controls. We conclude that internal confidence scores are not inherently calibrated to binding specificity and require validation against realistic decoys. To enable community benchmarking and method development, we release ~1.8 million AI-generated complex structures and guidance for the benchmarks ahead. Nomenclature We used 106 VHH-antigen complexes from the Protein Data Bank that have diverse structural and sequence properties. Their sequences were used for structure prediction using different AI models: Experimental complexes are the original 106 VHH-antigen complexes directly obtained from experimentally resolved structures from PDB Real complexes are structure-predicted models of the same 106 experimentally cognate VHH-antigen sequence pairs (i.e., each VHH paired with its true antigen) Shuffled complexes are structure-predicted models of artificially shuffled, non-cognate VHH-antigen sequence pairs (i.e., VHHs and antigens combined from different experimental complexes) Sequence: experimental = real ≠ shuffledStructure: experimental ≠ real ≠ shuffled Structural data Each structural file (.pdb) corresponds to one system (one VHH-antigen complex). The file name follows the format "system_i_j_PDBvhh_PDBantigen", where the numbers (i, j) correspond to row indices in the table with complexes (Supplementary Table 1, one row per experimental VHH-antigen complex with associated metadata), and the PDB IDs correspond to the source structures of the VHH and antigen. If the first and second indices/PDB IDs are identical (e.g., system_77_77_8H3X_8H3X), the complex represents the real VHH-antigen pair. If they differ (e.g., system_77_26_8H3X_6ZE1), the complex is shuffled, meaning the VHH and antigen were taken from different VHH-antigen complexes. In this case, the first index/PDB ID indicates the source of the VHH, and the second index/PDB ID indicates the source of the antigen. Preprint: https://www.biorxiv.org/content/10.64898/2026.03.02.709004v1Supplementary tables: https://github.com/csi-greifflab/ab_ag_champloo/tree/main/supplementary_tables



