yk0/proteinbase_interactions
收藏资源简介:
该数据集名为ProteinBase Interactions,专注于蛋白质结合剂与目标蛋白质之间的相互作用。每一行数据代表一个结合剂-目标对,包含结合剂的ProteinBase标识符、序列、目标蛋白质名称、序列、实验结合强度标签(如无、弱、中、强)、二元分类标签(0表示无结合,1表示有结合)、设计类别元数据(如scFv、小蛋白、肽段)以及是否为抗体衍生设计的标志(如纳米抗体或scFv)。数据集经过过滤,仅保留表达的结合剂,并对非抗体结合剂施加了最大序列相似性阈值(≤50%),而抗体衍生设计则不受此限制。结合亲和力以KD值的最小值表示,若无KD测量则为空。数据集来源于多个ProteinBase集合,适用于蛋白质-蛋白质相互作用分析、结合剂设计和抗体研究。
The dataset named ProteinBase Interactions focuses on interactions between protein binders and target proteins. Each row represents a single binder-target pair, including the ProteinBase binder identifier, binder sequence, target protein name, target sequence, experimental binding-strength labels (None, Weak, Medium, Strong), a binary classification label (0 for None, 1 for Weak, Medium, or Strong), design-class metadata (e.g., scFv, miniprotein, peptide), and an antibody flag indicating if the binder is antibody-derived (Nanobody or scFv). Rows are filtered to retain only expressed binders, with additional similarity filtering applied to non-antibody binders (max similarity ≤50%), while antibody binders are retained regardless. Binding affinity is represented by the minimum KD value, if available. The dataset is constructed from multiple ProteinBase collections and is suitable for protein-protein interaction studies, binder design, and antibody research.



