Pairing Heavy and Light Chains of Antibodies using Cross Attention
收藏资源简介:
Note to user The files have been organized according to the directory structure that is needed for running the scripts. This dataset contains the following files : 1. PHeLiX-main/ PHeLiX GitHub repository files. 2. PHeLiX-main/dataset/ Preprocessed_datasets/Cleaned, filtered, and standardized versions of the original OAS and PairedAbNGS datasets used for downstream analyses. Combined_Data/FASTA file of all unique VH–VL sequence pairs obtained by merging the OAS and PairedAbNGS datasets. VHVL_bindings/Pickle (.pkl) dictionary mapping each VH sequence to the list of experimentally observed paired VL sequences. The VH and VL sequences correspond to the unique VH-VL pairs above ( Combined_Data/ ) VH_VL_Separated/FASTA files (uniq_onlyVH.fasta and uniq_onlyVL.fasta) of unique VH and VL sequences extracted from the combined dataset. onlyVH_clustering_75/CD-HIT clustering results for VH sequences at 75% sequence identity. onlyVL_clustering_75/CD-HIT clustering results for VL sequences at 75% sequence identity. graph_at_75/NetworkX graph object formed using VH and VL clustering at 75% sequence identity. This graph contains only the positive VH-VL edges. MODEL_OOD_SPLIT/FASTA files defining the Train/Validation/In-Distribution (MODEL) and Out-of-Distribution (OOD) splits for VH sequences, VL sequences, and paired VH–VL sequences. OOD_onlyVH_clustering_90/CD-HIT clustering results for OOD VH sequences at 90% sequence identity. OOD_onlyVL_clustering_90/CD-HIT clustering results for OOD VL sequences at 90% sequence identity. MODEL_onlyVH_clustering_90/CD-HIT clustering results for MODEL VH sequences at 90% sequence identity. MODEL_onlyVL_clustering_90/CD-HIT clustering results for MODEL VL sequences at 90% sequence identity. graph_based_negatives_OOD/Graph-based positive and negative pair generation outputs for the OOD dataset: STATS.pkl: NetworkX graph object corresponding to the OOD set, along with other graph statistics like node degree, unique edge counts, etc. MISSING_0.9.pkl: Missing (negative) graph edges. POSITIVES_0.9.fasta: Positive VH–VL pairs. NEGATIVES_0.9.fasta: Graph-sampled negative VH–VL pairs. graph_based_negatives_MODEL/Contains graph-based positive and negative pair generation outputs for the MODEL Train/Validation/In-Distribution dataset: STATS.pkl: NetworkX graph object corresponding to the MODEL set, along with other graph statistics like node degree, unique edge counts, etc. MISSING_0.9.pkl: Missing (negative) graph edges. POSITIVES_0.9.fasta: Positive VH–VL pairs. NEGATIVES_0.9.fasta: Graph-sampled negative VH–VL pairs. Germline_info/Pickle (.pkl) files with germline annotations for VH–VL cluster-based graph formed for MODEL set, at 90% sequence identity threshold. 3. External_Benchmarking_Datasets/ MAGE_RBD/ VH-VL pairs generated by MAGE, including the 969 filtered pairs and 20 selected pairs for experimental validation. SynAbLib/ Cleaned SynAbLib dataset, with 1,198,151 VH-VL sequence pairs.



