claude-protein-binder-design
收藏资源简介:
该数据集(Claude蛋白结合子设计,数据发布v1.0)包含1,440个从头设计的微小蛋白结合子(长度为50-120个残基),靶向16个不同的蛋白质靶点。这些结合子由两个Claude模型(Mythos Preview生成900个设计,Opus 4.8生成540个设计)作为自主蛋白质设计代理完成。设计在两家合同研究组织(CRO)进行表征:Adaptyv Bio(无细胞表达,采用固定化设计的SPR/BLI动力学)和Twist Bioscience(Fc融合表达,采用六点抗原滴定的捕获SPR)。每个设计附有:两个供应商的结合性判定和动力学数据,包含原始传感图和报告图像;两个供应商的对比及最终评估结果;设计模型;来自十个结构预测器、每个种子最优的共折叠预测及评分;以及逐步设计来源信息。数据集以Parquet表形式组织,共20个配置子集(如design_summary、wetlab_summary、adaptyv_results、twist_fits、insilico_cofold_predictions等),每个设计还包含结构文件、传感器图PNG和曲线CSV文件。靶点包括:15-PGDH、BBF-14、BHRF1、SpCas9、EGFR、IL-7Rα、前肌生长抑制素(pro-myostatin)、成熟GDF-8、MBP、尼帕病毒G蛋白、PD-L1、RBX1、TNF-α、TREM2、TrkA、VEGF-A。其中对成熟GDF-8的120个设计湿实验测量结果因抗原聚集而无法使用,仅提供设计模型和预测数据;其余15个靶点的1,320个设计中,根据两个供应商的评估,354个被判定为结合子。该数据集适用于蛋白质设计、结合子预测、结构预测基准测试、以及蛋白质-蛋白质相互作用研究等任务。
This dataset (Claude Protein Binder Design, Data Release v1.0) contains 1,440 de novo designed small protein binders (50-120 residues long) targeting 16 different protein targets. These binders were designed by two Claude models (Mythos Preview generated 900 designs, Opus 4.8 generated 540 designs) as autonomous protein design agents. The designs were characterized at two contract research organizations (CROs): Adaptyv Bio (cell-free expression, SPR/BLI kinetics with immobilized designs) and Twist Bioscience (Fc-fusion expression, capture SPR with six-point antigen titration). Each design is accompanied by: binding status and kinetic data from both vendors, including raw sensorgrams and report images; comparison and final evaluation results from both vendors; design model; cofold predictions and scores from ten structure predictors for each best seed; and step-by-step design provenance. The dataset is organized as a Parquet table with 20 configuration subsets (e.g., design_summary, wetlab_summary, adaptyv_results, twist_fits, insilico_cofold_predictions, etc.), and each design also includes structure files, sensorgram PNGs, and curve CSV files. Targets include: 15-PGDH, BBF-14, BHRF1, SpCas9, EGFR, IL-7Rα, pro-myostatin, mature GDF-8, MBP, Nipah virus G protein, PD-L1, RBX1, TNF-α, TREM2, TrkA, VEGF-A. Among these, wet-lab measurements for 120 designs targeting mature GDF-8 were unusable due to antigen aggregation, so only design models and prediction data are provided. Of the remaining 1,320 designs for 15 targets, 354 were deemed binders based on both vendors evaluations. This dataset is suitable for tasks such as protein design, binder prediction, structure prediction benchmarking, and protein-protein interaction studies.
数据集概述
Claude protein binder design 是 Anthropic 发布的一个蛋白质设计数据集(v1.0),包含 1,440 个由 AI 设计的从头(de novo)微型蛋白结合剂,靶向 16 个目标蛋白,由两个 Claude 模型作为自主蛋白质设计智能体生成:
- Mythos Preview:900 个设计
- Opus 4.8:540 个设计
蛋白结合剂长度为 50 至 120 个氨基酸残基,实验表征由两家合同研究机构完成:
- Adaptyv Bio:无细胞表达,SPR/BLI 动力学检测(固定设计蛋白)
- Twist Bioscience:Fc 融合表达,六点抗原滴定的捕获 SPR
目标蛋白列表
15-PGDH、BBF-14、BHRF1、SpCas9、EGFR、IL-7Rα、潜在 GDF-8(pro-myostatin)、成熟 GDF-8、MBP、尼帕病毒 G、PD-L1、TREM2、TrkA、VEGF-A。
关键结果
- 成熟 GDF-8 的 120 个设计因抗原聚集和非特异性结合导致湿实验数据无效,未纳入(仅提供设计模型、共折叠和溯源信息)
- 其余 15 个靶点的 1,320 个设计中,354 个被两个供应商一致评定为结合剂
数据内容
数据集整合了每个设计的以下信息:
- 结合判定和两家供应商的动力学数据(含原始传感图和报告图像)
- 两个供应商的逐设计比较与最终评定
- 设计所用模型
- 十个结构预测器的种子最优共折叠预测(含逐种子评分)
- 逐步设计溯源信息
数据组织
- 20 个 Parquet 表:设计汇总、湿实验汇总与测量、两供应商结果/重复/读取/拟合曲线、表达与滴度、共折叠预测、表位接触、来源信息、靶点构建体、列字典等
- 逐设计文件夹:包含设计模型、共折叠结构、传感图 PNG、每曲线 CSV 等文件
- 配套文件:文档、清单(manifest)、提示词(prompts)、结构文件(structure_and_pae,含 mmCIF 格式和 PAE 矩阵)
- 所有表通过
uuid字段关联,full_name命名逐设计文件夹
许可证与引用
- 数据和文档:CC BY 4.0
- 脚本:MIT 许可
- 第三方材料保留各自的许可条款
- 引用方式见
CITATION.cff




