遇见数据集

viral-data-safety/tier1_excluded

收藏
Hugging Face2026-05-18 更新2026-04-26 收录
官方服务:

资源简介:

该数据集包含生物序列及其相关元数据。每个样本包括序列ID、描述、序列本身、长度、公共分类学ID、成员分类学ID列表、成员登录号列表,以及多个排除标志(tier1至tier6)。数据集分为训练集(约5764万样本)和验证集(约29万样本),总大小约32GB。

This dataset contains biological sequences with associated metadata. Each sample includes sequence ID, description, sequence, length, common taxid, list of member taxids, list of member accessions, and multiple exclusion flags (tier1 to tier6). The dataset is split into training (about 57.6 million samples) and validation (about 291k samples) sets, with a total size of ~32GB.

提供机构:
viral-data-safety
二维码
社区交流群
二维码
科研交流群
商业服务