Mgnify
收藏资源简介:
MGnify蛋白质目录数据集是MGnify蛋白质簇目录(包含mgy_clusters和mgy_proteins分区)的标准化FASTA分片集合,源自EMBL-EBI的MGnify微生物组分析资源,经过MegaData后下载流水线处理并标准化。数据集包含两个主要部分:一是sequences/目录下的FASTA格式蛋白质序列分片(使用Zstandard压缩),按源文件组织;二是tables/目录下的JSONL格式标准化元数据表。数据规模包括26个源文件、3,148个分片,压缩分片总大小为319.27 GiB,以及28个标准化表文件,压缩表总大小为2.17 TiB。重要注意事项:每个源目录下的shard-000000.fasta.zst仅为100条记录的验证样本,不包含在正式数据集中;metadata/和manifests/目录未提供,需通过流式处理分片重新生成统计信息。该数据集适用于宏基因组学、蛋白质序列分析、大规模生物信息学流水线等任务,需注意数据使用需遵循CC BY 4.0许可证并引用相关文献。
The MGnify protein catalog dataset is a standardized FASTA sharded collection of the MGnify protein cluster catalog (including the mgy_clusters and mgy_proteins partitions). It is derived from the MGnify microbiome analysis resource at EMBL-EBI, processed and standardized through the MegaData post-download pipeline. The dataset consists of two main components: first, FASTA-formatted protein sequence shards (compressed with Zstandard) in the sequences/ directory, organized by source files; second, standardized metadata tables in JSONL format in the tables/ directory. The data scale includes 26 source files, 3,148 shards with a total compressed size of 319.27 GiB, and 28 standardized table files with a total compressed size of 2.17 TiB. Important notes: The shard-000000.fasta.zst in each source directory is only a validation sample of 100 records and is not included in the official dataset; the metadata/ and manifests/ directories are not provided and need to be regenerated through streaming processing of shards for statistical information. The dataset is suitable for tasks such as metagenomics, protein sequence analysis, and large-scale bioinformatics pipelines, and users must comply with the CC BY 4.0 license and cite relevant literature.




