pulmo/ncbi-genbank-complete
收藏资源简介:
NCBI GenBank Complete 数据集是一个全面的基因组序列数据库,源自美国国立卫生研究院(NIH)的GenBank®,这是一个包含所有公开可用DNA序列的注释集合。GenBank 是国际核苷酸序列数据库协作(INSDC)的一部分,该协作还包括日本DNA数据库(DDBJ)和欧洲核苷酸档案(ENA),这三个组织每日交换数据。数据集已处理为适合机器学习训练的parquet格式,包含基因组序列及其对应的登录号(accession)。需要注意的是,与RefSeq数据库相比,GenBank 是冗余的,因为它包含作者提交的原始序列,未经过去重处理。数据以核苷酸字符串形式表示,包括A、C、G、T和N等字符,代表DNA或RNA序列。该数据集可用于训练大规模基因组基础模型、进行广泛的序列分类或研究遗传多样性。
The NCBI GenBank Complete dataset is a comprehensive genomic sequence database derived from GenBank® of the U.S. National Institutes of Health (NIH), which is an annotated collection of all publicly available DNA sequences. GenBank is part of the International Nucleotide Sequence Database Collaboration (INSDC), which also includes the DNA Data Bank of Japan (DDBJ) and the European Nucleotide Archive (ENA); these three organizations exchange data on a daily basis. The dataset has been processed into Parquet format suitable for machine learning training, containing genomic sequences and their corresponding accession numbers. Notably, compared with the RefSeq database, GenBank is redundant, as it contains the original sequences submitted by authors without deduplication processing. The data is represented as nucleotide strings, including characters such as A, C, G, T, and N that represent DNA or RNA sequences. This dataset can be used for training large-scale genomic foundation models, conducting extensive sequence classification, or researching genetic diversity.



