Diamond2GO Database – Version 2025-08-12 (nr_clean_d2go)
收藏资源简介:
This is the updated Diamond2GO reference database built on 12th August 2025. It is a DIAMOND-formatted protein database (`.dmnd`) consisting of over 27 million sequences derived from the NCBI `nr` dataset, filtered to include only those with Gene Ontology (GO) annotations, and redundancy reduction using MMseqs2 (95% similarity). This version improves sensitivity and annotation coverage compared to the original 2023 release used in the published D2GO manuscript, and the earlier 2025 release. This database is intended for use with the Diamond2GO tool, which enables rapid GO-term annotation and enrichment analysis for high-throughput sequencing datasets. For reproducibility of results published using the earlier version (699,409 sequences), please refer to the [v1.0.0 release] https://github.com/rhysf/Diamond2GO/releases/tag/6a035ce
本数据集为2025年8月12日更新的Diamond2GO参考数据库。 它是一款采用DIAMOND格式构建的蛋白质数据库(文件后缀为.dmnd),序列源自NCBI的nr数据集,总计包含超过2700万条蛋白质序列;筛选保留了所有带有基因本体(Gene Ontology,GO)注释的序列,并通过MMseqs2工具以95%相似度阈值完成了冗余序列去除。 相较于2023年发表D2GO论文中使用的初始版本以及2025年的早期发布版本,本版本的检测灵敏度与注释覆盖度均得到了提升。 本数据库专为Diamond2GO工具设计,可用于高通量测序数据集的快速GO功能注释与富集分析。 若需复现使用早期版本(包含699,409条序列)所发表的研究结果,请参阅[v1.0.0版本发布页] https://github.com/rhysf/Diamond2GO/releases/tag/6a035ce



