遇见数据集

Protein Set Transformer: A protein-based genome language model to power high diversity viromics

收藏
DataONE2026-04-19 更新2026-05-19 收录
官方服务:

资源简介:

Exponential increases in microbial and viral genomic data demand transformational advances in scalable, generalizable frameworks for their interpretation. Standard homology-based functional analyses are hindered by the rapid divergence of microbial and especially viral genomes and proteins that significantly decreases the volume of usable data. Here, we present Protein Set Transformer (PST), a protein-based genome language model that models genomes as sets of proteins without considering sparsely available functional labels. Trained on >100k viruses, PST outperformed other homology- and language model-based approaches for relating viral genomes based on shared protein content. Further, PST demonstrated protein structural and functional awareness by clustering capsid-fold-containing proteins with known capsid proteins and uniquely clustering late gene proteins within related viruses. Our data establish PST as a valuable method for diverse viral genomics, ecology, and evolutionary appl..., Genomes used to create this dataset are publicly available, and all data within this dataset were generated by the study. See the manuscript for details., , # Data from: Protein Set Transformer: A protein-based genome language model to power high diversity viromics These datasets are associated with the manuscript \"Protein Set Transformer: A protein-based genome language model powers high diversity viromics\". ([https://doi.org/10.1101/2024.07.26.605391](https://doi.org/10.1101/2024.07.26.605391)) ## Dataset descriptions This statement is relevant for the file references in the manuscript code: We refer to the processed data directly used for figure making as \"**Source Data**\". Major datasets that are used throughout the manuscript are referred to as \"**Datasets**\". Due to the size of the \"datasets\", we have provided each dataset folder as individual tarballs. \"Supplementary Data\" and \"Supplementary Tables\" are manuscript-associated tables, with \"Supplementary Data\" being too large otherwise be manuscript tables. ### Main data | File | Description ..., , **Changes after Sep 19, 2024:** * Added * `foldseek_databases.tar.gz` * Precomputed foldseek 3Di databases for each test dataset * `PST-TL-P__small.ckpt.gz` * New pretrained model checkpoint for model trained with triplet loss and tuned with protein diversity groups * `PST-TL-P__large.ckpt.gz` * New pretrained model checkpoint for model trained with triplet loss and tuned with protein diversity groups * `PST-MLM.tar.gz` * New pretrained model checkpoints for models trained with masked language modeling loss * `PST_training_set_PST-TL-P__large_protein_embeddings.h5` * `IMGVRv4_test_set_PST-TL-P__large_protein_embeddings.h5` * `MGnify_set_PST-TL-P__large_protein_embeddings.h5` * `PST_training_set_PST-TL-P__small_protein_embeddings.h5` * `IMGVRv4_test_set_PST-TL-P__small_protein_embeddings.h5` * `MGnify_set_PST-TL-P__small_protein_embeddings.h5` * Changed * `esm-large_protein_embeddings.tar.gz` * Now part of `esm_embeddings.tar.gz` * Includes...

微生物与病毒基因组数据呈指数级增长,这要求我们在可扩展、可泛化的解读框架上取得变革性进展。传统基于同源性的功能分析受限于微生物尤其是病毒基因组与蛋白质的快速演化分化,大幅缩减了可用数据的体量。在此,我们提出**蛋白质集Transformer(Protein Set Transformer, PST)**,这是一种基于蛋白质的基因组语言模型,将基因组建模为蛋白质集合,无需考虑稀缺可用的功能标签。在超过10万条病毒序列上训练后,PST在基于共享蛋白质内容关联病毒基因组的任务中,表现优于其他基于同源性与语言模型的方法。此外,PST展现出对蛋白质结构与功能的感知能力:它将包含衣壳折叠结构的蛋白质与已知衣壳蛋白质聚为一类,并在相关病毒中单独将晚期基因蛋白质聚类。本研究数据证实,PST可作为适用于多样化病毒基因组学、生态学与进化研究的可靠工具。用于构建本数据集的基因组均为公开可用,且本数据集内的所有数据均由本研究生成,详细信息请参阅论文手稿。 # 数据来源:《蛋白质集Transformer:赋能高多样性病毒组学的基于蛋白质的基因组语言模型》 本数据集与论文《蛋白质集Transformer:赋能高多样性病毒组学的基于蛋白质的基因组语言模型》相关联,DOI链接:https://doi.org/10.1101/2024.07.26.605391 ## 数据集说明 本说明适配论文代码中的文件引用规则:我们将直接用于绘制图表的处理后数据定义为**源数据(Source Data)**;将全文核心使用的数据集称为**数据集(Datasets)**。鉴于数据集整体体量庞大,我们将每个数据集文件夹打包为独立的tar压缩包。“补充数据(Supplementary Data)”与“补充表格(Supplementary Tables)”均为论文配套附表,其中“补充数据”因体积过大,无法以论文附表形式呈现。 ### 核心数据集 | 文件名称 | 描述 ..., , **2024年9月19日后的更新内容:** * 新增文件 * `foldseek_databases.tar.gz`:针对各测试数据集预先计算的Foldseek 3Di数据库 * `PST-TL-P__small.ckpt.gz`:基于三元组损失训练、并通过蛋白质多样性分组微调的小型预训练模型检查点文件 * `PST-TL-P__large.ckpt.gz`:基于三元组损失训练、并通过蛋白质多样性分组微调的大型预训练模型检查点文件 * `PST-MLM.tar.gz`:基于掩码语言建模损失训练的预训练模型检查点合集 * `PST_training_set_PST-TL-P__large_protein_embeddings.h5`:PST训练集的大型PST-TL-P模型蛋白质嵌入文件 * `IMGVRv4_test_set_PST-TL-P__large_protein_embeddings.h5`:IMGVRv4测试集的大型PST-TL-P模型蛋白质嵌入文件 * `MGnify_set_PST-TL-P__large_protein_embeddings.h5`:MGnify数据集的大型PST-TL-P模型蛋白质嵌入文件 * `PST_training_set_PST-TL-P__small_protein_embeddings.h5`:PST训练集的小型PST-TL-P模型蛋白质嵌入文件 * `IMGVRv4_test_set_PST-TL-P__small_protein_embeddings.h5`:IMGVRv4测试集的小型PST-TL-P模型蛋白质嵌入文件 * `MGnify_set_PST-TL-P__small_protein_embeddings.h5`:MGnify数据集的小型PST-TL-P模型蛋白质嵌入文件 * 变更文件 * `esm-large_protein_embeddings.tar.gz`:现已整合至`esm_embeddings.tar.gz`,包含内容详见原说明。

创建时间:
2026-04-20
二维码
社区交流群
二维码
科研交流群
商业服务