Awesome-Large-Biology-Models
收藏资源简介:
该仓库是一个关于大型生物学模型的数据集合集,专注于收集、整理和索引与DNA、RNA、蛋白质、单细胞和生物文本等生物学领域相关的多个数据集资源,包括通用数据集和特定领域数据集,以支持大型生物学模型的研究和开发。
This repository is a curated collection of datasets for large biological models. It focuses on collecting, organizing, and indexing multiple dataset resources related to biological fields including DNA, RNA, proteins, single cells, biological text, etc., covering both general-purpose and domain-specific datasets to support the research and development of large biological models.
Awesome Large Biology Models 数据集详情
概述
这是一个专注于生物学领域大模型的资源汇总仓库,涵盖DNA、RNA、蛋白质、单细胞和生物文本等多种生物数据类型的模型与数据集。仓库主要分为“模型与代码”和“数据集”两大部分,旨在为生物学研究人员和开发者提供结构化的资源索引。
模型与代码
DNA 模型
收录了10个DNA序列分析模型,发布时间从2021年至2023年:
| 模型名称 | 发表平台/时间 | 代码/资源 |
|---|---|---|
| Enformer | Nature Methods, 2021-10 | 代码 |
| DNABERT | Bioinformatics, 2021-02 | 代码 |
| GenSLMs | BioRxiv, 2022-10 | - |
| Nucleotide Transformer | BioRxiv, 2023-01 | 代码 |
| DNABERT-2 | Arxiv, 2023-06 | 代码 |
| HyenaDNA | Arxiv, 2023-06 | 代码 |
| DNAGPT | Arxiv, 2023-07 | 代码 |
| EpiGePT | BioRxiv, 2023-07 | 网站 |
| Predicting RNA-seq coverage | BioRxiv, 2023-09 | - |
| ChromTransfer | NAR Genomics, 2023-03 | 代码 |
RNA 模型
收录了7个RNA相关模型,发布时间从2022年至2023年,涵盖RNA结构预测、剪接预测、mRNA设计等任务:
- RNA-FM (2022-08): 非编码RNA/UTR功能预测,代码
- SpliceBERT (2023-02): 前体mRNA剪接预测,代码
- UNI-RNA (2023-07): 通用RNA预训练模型
- CodonBERT (2023-09): mRNA设计与优化
- DRfold (2023-09): RNA结构预测,代码
- An RNA foundation model (2023-09)
- Design of prime-editing guide RNAs (2023-10)
蛋白质模型
收录了14个蛋白质模型,时间跨度从2022年至2023年,包括序列模型和结构模型:
- ProtTrans (TPAMI, 2022-10): 序列模型,代码
- ProtGPT2 (Nature Communications, 2022-07): 蛋白质设计
- ProGen2 (Arxiv, 2022-06): 序列模型,代码
- OntoProtein (ICLR, 2022-06): 基因本体嵌入,代码
- AlphaFold-Multimer (BioRxiv, 2022-03): 蛋白质复合物结构预测,代码
- xTrimoPGLM (Arxiv, 2022-07): 统一100B参数预训练模型
- ESM (Science, 2023-03): 进化尺度蛋白质结构预测,代码
- RFdiffusion (Nature, 2023-07): 蛋白质结构从头设计,代码
- ESM-variants (Nature Genetics, 2023-08): 疾病变异效应预测,代码
- AlphaMissense (Science, 2023-09): 错义变异效应预测,代码
- Latent Diffusion Model (Arxiv, 2023-05)
- Large language models generate functional protein sequences (Nature Biotechnology, 2023-01)
- Large language models improve annotation of viral proteins (Preprint, 2023-05)
- Exploring the Protein Sequence Space (CSH Perspectives, 2023-10)
- Atom-by-atom protein generation (Arxiv, 2023-08)
- SaLT&PepPr (Communications Biology, 2023-10)
单细胞模型
收录了4个单细胞分析模型:
- scGPT (BioRxiv/Arxiv, 2023-04/09): 单细胞多组学基础模型,代码
- scBERT (Nature Machine Intelligence, 2023-09): 细胞类型注释,代码
- Revolutionizing Single Cell Analysis (Arxiv, 2023-04)
生物文本模型
收录了4个生物医学文本处理模型:
- BioBERT (Bioinformatics, 2019-09): 生物医学文本挖掘,代码
- BioGPT (Briefings in Bioinformatics, 2022-12): 生物医学文本生成与挖掘,代码
- Large language models in biomedical NLP (Arxiv, 2022-05): 基准测试与基线,代码
- GeneGPT (Preprint, 2023-05): 增强大语言模型访问生物信息,代码
数据集
通用数据集
6个主要生物数据库,涵盖DNA/RNA/蛋白质多种数据类型:
- NCBI Database: 链接
- Ensembl Database: 链接
- UCSC Database: 链接
- RCSB PDB Database: 蛋白质结构数据,链接
- Uniprot Database: 蛋白质序列数据,链接
- STRING Database: 蛋白质相互作用数据,链接
DNA 数据集
- Genome Understanding Evaluation (GUE): 序列分类数据集,用于DNABERT-2模型基准测试,链接
RNA 数据集
- Genotype-Tissue Expression (GTEx): 人类组织基因调控效应数据,链接
- Master database of All possible RNA sequences (MARS): 所有可能RNA序列数据库,链接
蛋白质数据集
- MGnify: 微生物组序列数据,链接
- UniRef: UniProt参考聚类,链接
- AlphaFoldDB: DeepMind AlphaFold产生的蛋白质结构数据,链接
- SidechainNet: 全原子蛋白质结构数据集,链接
- Colabfold DB: 蛋白质序列数据集,链接
- Big Fantastic Dataset (BFD): 通用蛋白质数据集,链接
单细胞数据集
待补充(TBD)
生物文本数据集
待补充(TBD)




