five

Protein Set Transformer: A protein-based genome language model to power high diversity viromics

收藏
DataONE2025-06-18 更新2025-08-02 收录
下载链接:
https://search.dataone.org/view/sha256:4c4fd7bdae194c5655843265f4bf67af9bdea40972ce3daa1c893cb6e8e4b425
下载链接
链接失效反馈
官方服务:
资源简介:
Exponential increases in microbial and viral genomic data demand transformational advances in scalable, generalizable frameworks for their interpretation. Standard homology-based functional analyses are hindered by the rapid divergence of microbial and especially viral genomes and proteins that significantly decreases the volume of usable data. Here, we present Protein Set Transformer (PST), a protein-based genome language model that models genomes as sets of proteins without considering sparsely available functional labels. Trained on >100k viruses, PST outperformed other homology- and language model-based approaches for relating viral genomes based on shared protein content. Further, PST demonstrated protein structural and functional awareness by clustering capsid-fold-containing proteins with known capsid proteins and uniquely clustering late gene proteins within related viruses. Our data establish PST as a valuable method for diverse viral genomics, ecology, and evolutionary appl..., Genomes used to create this dataset are publicly available, and all data within this dataset were generated by the study. See the manuscript for details., , # Data from: Protein Set Transformer: A protein-based genome language model to power high diversity viromics These datasets are associated with the manuscript \"Protein Set Transformer: A protein-based genome language model powers high diversity viromics\". ([https://doi.org/10.1101/2024.07.26.605391](https://doi.org/10.1101/2024.07.26.605391)) ## Dataset descriptions This statement is relevant for the file references in the manuscript code: We refer to the processed data directly used for figure making as \"Supplementary **Data**\". Major datasets that are used throughout the manuscript are referred to as \"**Datasets**\". Due to the size of the \"datasets\", we have provided each dataset folder as individual tarballs. ### Main data | File | Description | |...,
创建时间:
2025-06-19
5,000+
优质数据集
54 个
任务类型
进入经典数据集
二维码
社区交流群

面向社区/商业的数据集话题

二维码
科研交流群

面向高校/科研机构的开源数据集话题

数据驱动未来

携手共赢发展

商业合作