遇见数据集

dynotx/synthetic_dimers

收藏
Hugging Face2026-03-18 更新2026-03-29 收录
官方服务:

资源简介:

--- license: cc-by-4.0 --- # Dyno Synthetic Dimer Dataset The datasets in this repo were used to train Dyno Psi-0 and Dyno Psi-1. This dataset was curated using predicted protein structures from the **AlphaFold Protein Structure Database**, subset to entries from **AFDB50**. Dimers were curated by using domain annotations from the **TED domain annotations provided by the CATH database** to identify domains within the same monomer structure that share an interface. These were further clustered using **Foldseek** to produce cluster representatives. A further refined set of dimers was identified by refolding all the cluster representatives with **AlphaFold2**, filtering by interface metrics, and refined using **PyRosetta FastRelax**. For more information on how this dataset was generated and used, please see our [white paper](https://dynopsi.dynotx.com/dynopsi_whitepaper.pdf). # Summary This dataset contains the three compressed archives along with an index file. Be aware that two of the archives contain a large number of files, and the full dataset takes up ~800GB of disk space. Consider only downloading the cluster representatives or the AF2-filtered set if you don't need the full set. The archives contain structure files for each synthetic dimer. Each dimer is assigned an ID based on the AFDB entry it came from along with the TED-annotated domains that constitute the chains of the synthetic dimer (e.g. `AF-A0A0B7FL75-F1-model_v4_TED01_TED02`). The structure files each have two protein chains, A and B, corresponding to these two domains. | Name | Description | Number of files | Uncompressed size | |------|-------------|-----------------|-------------------| | `cifs_cluster_reps.tar.gz` | MMCIF files for the cluster representatives | 1063207 | 198 GB | | `cifs_noncluster_reps.tar.gz` | MMCIF files for the other dimers that are not cluster reps | 3007752 | 605 GB | | `pdbs_filtered_relaxed.tar.gz` | PDB files for the cluster reps that pass additional AF2 refolding filters and were further relaxed with PyRosetta FastRelax | 32253 | 10 GB | `index_df.tsv` fields: | Field | Description | Example | |-------|-------------|---------| | `dimer_id` | `{afdb_id}_{domain_id1}\_{domain_id2}` | AF-A0A0B7FL75-F1-model_v4_TED01_TED02 | | `cluster_rep` | the dimer_id of the cluster representative for the cluster this dimer belongs to | AF-A0A0B7FL75-F1-model_v4_TED01_TED02 | | `is_cluster_rep` | whether or not this dimer is a cluster representative | True | | `afdb_id` | the AFDB ID of the original monomer | AF-A0A0B7FL75-F1-model_v4 | | `start_{1/2}` | the start index of the 1st/2nd domain in the original monomer, 1-indexed | 5 | | `end_{1/2}` | the end index of the 1st/2nd domain in the original monomer, 1-indexed | 145 | | `num_res_{1/2}` | the length of the 1st/2nd domain | 141 | | `num_interface_res_{1/2}` | the number of residues in the 1st/2nd domain for which the CA is within 10Å of a CA in the other domain | 14 | | `interface_idxs_{1/2}` | indices of the interface residues in the 1st/2nd domain, 1-indexed relative to the full-length monomer | 10,11,12,14,94,96,97,100,129,130,131,132,133,134 | | `total_length` | total length of the dimer | 370 | | `pass_af2_filter` | whether or not this dimer passed AlphaFold2 refolding interface confidence metrics | False | # Citations ```bibtex @article{varadi2021alphafold-f41, year = {2021}, title = {{AlphaFold} Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models}, author = {Varadi, Mihaly and Anyango, Stephen and Deshpande, Mandar and Nair, Sreenath and Natassia, Cindy and Yordanova, Galabina and Yuan, David and Stroe, Oana and Wood, Gemma and Laydon, Agata and Žídek, Augustin and Green, Tim and Tunyasuvunakool, Kathryn and Petersen, Stig and Jumper, John and Clancy, Ellen and Green, Richard and Vora, Ankur and Lutfi, Mira and Figurnov, Michael and Cowie, Andrew and Hobbs, Nicole and Kohli, Pushmeet and Kleywegt, Gerard and Birney, Ewan and Hassabis, Demis and Velankar, Sameer}, journal = {Nucleic Acids Research}, issn = {0305-1048}, doi = {10.1093/nar/gkab1061}, pmid = {34791371}, pmcid = {{PMC}8728224}, pages = {D439--D444}, number = {D1}, volume = {50} } @article{lau2024exploring-e2f, year = {2024}, title = {Exploring structural diversity across the protein universe with The Encyclopedia of Domains}, author = {Lau, Andy M. and Bordin, Nicola and Kandathil, Shaun M. and Sillitoe, Ian and Waman, Vaishali P. and Wells, Jude and Orengo, Christine A. and Jones, David T.}, journal = {Science}, issn = {0036-8075}, doi = {10.1126/science.adq4946}, pmid = {39480926}, pages = {eadq4946}, number = {6721}, volume = {386} } @article{kempen2022fast-33d, year = {2022}, title = {Fast and accurate protein structure search with Foldseek}, author = {Kempen, Michel van and Kim, Stephanie S and Tumescheit, Charlotte and Mirdita, Milot and Lee, Jeongjae and Gilchrist, Cameron L M and Söding, Johannes and Steinegger, Martin}, journal = {{bioRxiv}}, doi = {10.1101/2022.02.07.479398}, pages = {2022.02.07.479398} } @article{jumper2021highly-969, year = {2021}, title = {Highly accurate protein structure prediction with {AlphaFold}}, author = {Jumper, John and Evans, Richard and Pritzel, Alexander and Green, Tim and Figurnov, Michael and Ronneberger, Olaf and Tunyasuvunakool, Kathryn and Bates, Russ and Žídek, Augustin and Potapenko, Anna and Bridgland, Alex and Meyer, Clemens and Kohl, Simon A. A. and Ballard, Andrew J. and Cowie, Andrew and Romera-Paredes, Bernardino and Nikolov, Stanislav and Jain, Rishub and Adler, Jonas and Back, Trevor and Petersen, Stig and Reiman, David and Clancy, Ellen and Zielinski, Michal and Steinegger, Martin and Pacholska, Michalina and Berghammer, Tamas and Bodenstein, Sebastian and Silver, David and Vinyals, Oriol and Senior, Andrew W. and Kavukcuoglu, Koray and Kohli, Pushmeet and Hassabis, Demis}, journal = {Nature}, issn = {0028-0836}, doi = {10.1038/s41586-021-03819-2}, pmid = {34265844}, pmcid = {{PMC}8371605}, pages = {583--589}, number = {7873}, volume = {596} } @article{chaudhury2010pyrosetta-aec, year = {2010}, title = {{PyRosetta}: a script-based interface for implementing molecular modeling algorithms using Rosetta}, author = {Chaudhury, Sidhartha and Lyskov, Sergey and Gray, Jeffrey J.}, journal = {Bioinformatics}, issn = {1367-4803}, doi = {10.1093/bioinformatics/btq007}, pmid = {20061306}, pmcid = {{PMC}2828115}, pages = {689--691}, number = {5}, volume = {26} } ```

许可证:CC-BY-4.0 # Dyno合成二聚体数据集 本仓库中的数据集用于训练Dyno Psi-0与Dyno Psi-1。本数据集基于**AlphaFold蛋白质结构数据库(AlphaFold Protein Structure Database)**的预测蛋白质结构构建,子集取自**AFDB50**。通过使用**CATH数据库提供的TED结构域注释**识别同一单体结构内共享相互作用界面的结构域,从而筛选得到二聚体。随后使用**Foldseek**进行聚类,生成聚类代表序列。进一步通过**AlphaFold2**对所有聚类代表序列进行重新折叠,基于界面指标进行筛选,并使用**PyRosetta FastRelax**进行精细化优化。有关本数据集的生成与使用细节,请参阅我们的[白皮书](https://dynopsi.dynotx.com/dynopsi_whitepaper.pdf)。 # 数据集概览 本数据集包含三个压缩归档文件与一个索引文件。请注意,其中两个归档包含大量文件,完整数据集占用约800GB磁盘空间。若无需完整数据集,可仅下载聚类代表序列或经AlphaFold2筛选后的子集。 每个归档均包含各合成二聚体的结构文件。每个二聚体均根据其来源的AFDB条目以及构成合成二聚体链的TED注释结构域分配唯一ID(例如`AF-A0A0B7FL75-F1-model_v4_TED01_TED02`)。结构文件包含两条蛋白质链A与B,分别对应上述两个结构域。 | 文件名 | 描述 | 文件数量 | 未压缩大小 | |------|-------------|-----------------|-------------------| | `cifs_cluster_reps.tar.gz` | 聚类代表序列的MMCIF格式文件 | 1063207 | 198 GB | | `cifs_noncluster_reps.tar.gz` | 非聚类代表二聚体的MMCIF格式文件 | 3007752 | 605 GB | | `pdbs_filtered_relaxed.tar.gz` | 经额外AlphaFold2重折叠筛选、并通过PyRosetta FastRelax精细化优化的聚类代表序列的PDB格式文件 | 32253 | 10 GB | `index_df.tsv`各字段说明如下: | 字段名 | 描述 | 示例 | |-------|-------------|---------| | `dimer_id` | 格式为`{afdb_id}_{domain_id1}_{domain_id2}`的二聚体唯一标识符 | `AF-A0A0B7FL75-F1-model_v4_TED01_TED02` | | `cluster_rep` | 当前二聚体所属聚类的聚类代表二聚体ID | `AF-A0A0B7FL75-F1-model_v4_TED01_TED02` | | `is_cluster_rep` | 当前二聚体是否为聚类代表序列 | `True` | | `afdb_id` | 原始单体的AFDB ID | `AF-A0A0B7FL75-F1-model_v4` | | `start_{1/2}` | 原始单体中第1/2个结构域的起始索引(采用1起始计数) | `5` | | `end_{1/2}` | 原始单体中第1/2个结构域的终止索引(采用1起始计数) | `145` | | `num_res_{1/2}` | 第1/2个结构域的氨基酸残基总长度 | `141` | | `num_interface_res_{1/2}` | 第1/2个结构域中,其α碳原子与另一结构域的α碳原子间距小于10Å的残基数量 | `14` | | `interface_idxs_{1/2}` | 第1/2个结构域中界面残基的索引(相对于全长单体的1起始计数) | `10,11,12,14,94,96,97,100,129,130,131,132,133,134` | | `total_length` | 二聚体的总残基长度 | `370` | | `pass_af2_filter` | 当前二聚体是否通过AlphaFold2重折叠界面置信度筛选 | `False` | # 引用文献 bibtex @article{varadi2021alphafold-f41, year = {2021}, title = {AlphaFold蛋白质结构数据库:利用高精度模型大规模拓展蛋白质序列空间的结构覆盖范围}, author = {Varadi, Mihaly and Anyango, Stephen and Deshpande, Mandar and Nair, Sreenath and Natassia, Cindy and Yordanova, Galabina and Yuan, David and Stroe, Oana and Wood, Gemma and Laydon, Agata and Žídek, Augustin and Green, Tim and Tunyasuvunakool, Kathryn and Petersen, Stig and Jumper, John and Clancy, Ellen and Green, Richard and Vora, Ankur and Lutfi, Mira and Figurnov, Michael and Cowie, Andrew and Hobbs, Nicole and Kohli, Pushmeet and Kleywegt, Gerard and Birney, Ewan and Hassabis, Demis and Velankar, Sameer}, journal = {Nucleic Acids Research}, issn = {0305-1048}, doi = {10.1093/nar/gkab1061}, pmid = {34791371}, pmcid = {PMC8728224}, pages = {D439--D444}, number = {D1}, volume = {50} } @article{lau2024exploring-e2f, year = {2024}, title = {探索蛋白质宇宙的结构多样性:结构域百科全书}, author = {Lau, Andy M. and Bordin, Nicola and Kandathil, Shaun M. and Sillitoe, Ian and Waman, Vaishali P. and Wells, Jude and Orengo, Christine A. and Jones, David T.}, journal = {Science}, issn = {0036-8075}, doi = {10.1126/science.adq4946}, pmid = {39480926}, pages = {eadq4946}, number = {6721}, volume = {386} } @article{kempen2022fast-33d, year = {2022}, title = {利用Foldseek实现快速准确的蛋白质结构搜索}, author = {Kempen, Michel van and Kim, Stephanie S and Tumescheit, Charlotte and Mirdita, Milot and Lee, Jeongjae and Gilchrist, Cameron L M and Söding, Johannes and Steinegger, Martin}, journal = {bioRxiv}, doi = {10.1101/2022.02.07.479398}, pages = {2022.02.07.479398} } @article{jumper2021highly-969, year = {2021}, title = {利用AlphaFold实现高精度蛋白质结构预测}, author = {Jumper, John and Evans, Richard and Pritzel, Alexander and Green, Tim and Figurnov, Michael and Ronneberger, Olaf and Tunyasuvunakool, Kathryn and Bates, Russ and Žídek, Augustin and Potapenko, Anna and Bridgland, Alex and Meyer, Clemens and Kohl, Simon A. A. and Ballard, Andrew J. and Cowie, Andrew and Romera-Paredes, Bernardino and Nikolov, Stanislav and Jain, Rishub and Adler, Jonas and Back, Trevor and Petersen, Stig and Reiman, David and Clancy, Ellen and Zielinski, Michal and Steinegger, Martin and Pacholska, Michalina and Berghammer, Tamas and Bodenstein, Sebastian and Silver, David and Vinyals, Oriol and Senior, Andrew W. and Kavukcuoglu, Koray and Kohli, Pushmeet and Hassabis, Demis}, journal = {Nature}, issn = {0028-0836}, doi = {10.1038/s41586-021-03819-2}, pmid = {34265844}, pmcid = {PMC8371605}, pages = {583--589}, number = {7873}, volume = {596} } @article{chaudhury2010pyrosetta-aec, year = {2010}, title = {PyRosetta:基于脚本的Rosetta分子建模算法实现接口}, author = {Chaudhury, Sidhartha and Lyskov, Sergey and Gray, Jeffrey J.}, journal = {Bioinformatics}, issn = {1367-4803}, doi = {10.1093/bioinformatics/btq007}, pmid = {20061306}, pmcid = {PMC2828115}, pages = {689--691}, number = {5}, volume = {26} }

提供机构:
dynotx
二维码
社区交流群
二维码
科研交流群
商业服务