遇见数据集

LiteFold/Evolutionary

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

该数据集名为进化MSA数据,是一个预计算的进化序列比对数据存档,采用便于在HuggingFace Hub上托管和下载的存档格式。数据集旨在为需要现成多序列比对(MSA)或缓存文件的工作流提供支持,用户无需从原始序列数据库重新构建这些文件。数据内容包括msa_cache/和mmseqs30/等目录,包含大量.npz、.fasta、.parquet等文件类型,总存档大小约为19.72 GiB。数据集通过metadata.csv提供可搜索的索引,保留原始文件路径,并支持通过Python API进行浏览和提取操作。适用于生物学、蛋白质研究、序列对齐和特征提取等任务。

This dataset, named Evolutionary MSA Data, is a precomputed evolutionary sequence alignment data archive adopting an archive format optimized for hosting and downloading on the HuggingFace Hub. It is designed to support workflows requiring ready-to-use multiple sequence alignment (MSA) or cache files, sparing users the need to rebuild these files from raw sequence databases. The dataset includes directories such as msa_cache/ and mmseqs30/, which contain a large number of files in formats including .npz, .fasta, .parquet and other types, with a total archive size of approximately 19.72 GiB. A searchable index is provided via metadata.csv, which retains the original file paths, and the dataset supports browsing and extraction operations through the Python API. It is applicable to tasks such as biological research, protein studies, sequence alignment and feature extraction.

提供机构:
LiteFold
二维码
社区交流群
二维码
科研交流群
商业服务