遇见数据集

binchenlab/GEOMeta

收藏
Hugging Face2026-05-13 更新2026-05-31 收录
官方服务:

资源简介:

GEOMeta是一个大规模的人类批量RNA-seq数据集,源自ARCHS4资源,包含约474,000个样本的转录本丰度谱,涵盖训练集、测试集和保留测试集。每个样本都标注了标准化的元数据,包括性别、器官系统、疾病类别、年龄组和实验设置。数据集构建过程包括从ARCHS4的HDF5文件中提取转录本丰度谱,排除有问题的基因索引范围,提取基因标识符(如Ensembl ID和基因符号)和样本注释(如GSM编号、样本名称和系列ID),并组装成AnnData对象。然后通过GSM编号合并精心整理的元数据表,去除重复的GSM条目。数据集文件包括训练、测试和保留测试集的AnnData文件(.h5ad格式)和CSV元数据文件,以及用于生成AnnData文件的Python脚本。AnnData结构包括样本级元数据(如性别、器官系统、疾病、年龄组和实验设置)、基因级元数据(如Ensembl ID和基因符号)和原始转录本丰度矩阵(稀疏格式,覆盖65,186个基因)。元数据CSV列包括GSE系列编号、GSM样本编号、提交年份、提交国家、测序平台、RNA库类型、实验设置、是否涉及扰动、疾病标签、器官、性别和年龄组。数据集可用于生物信息学和机器学习研究,支持转录组分析和元数据挖掘。

GEOMeta is a large-scale human bulk RNA-seq dataset derived from the ARCHS4 resource, containing transcript abundance profiles for approximately 474,000 samples across training, test, and held-out test splits. Each sample is annotated with standardized metadata including sex, organ system, disease category, age group, and experimental setting. The dataset construction involves extracting transcript abundance profiles from the ARCHS4 human reference HDF5 file, excluding problematic gene index ranges, and assembling gene identifiers (e.g., Ensembl ID and gene symbol) and sample annotations (e.g., GSM accession, sample name, and series ID) into an AnnData object. Curated metadata tables are then merged by GSM accession, with duplicate GSM entries removed. The dataset files include AnnData files (.h5ad format) and CSV metadata files for training, test, and held-out test splits, along with a Python script for generating AnnData files. The AnnData structure comprises sample-level metadata (e.g., gender, organ system, disease, age group, and experimental setting), gene-level metadata (e.g., Ensembl ID and gene symbol), and a raw transcript abundance matrix in sparse format covering 65,186 genes. The metadata CSV columns include GSE series accession, GSM sample accession, submission year, submitting country, sequencing platform, RNA library type, experimental setting, perturbation involvement, disease label, organ, sex, and age group. The dataset is suitable for bioinformatics and machine learning research, enabling transcriptome analysis and metadata exploration.

提供机构:
binchenlab
二维码
社区交流群
二维码
科研交流群
商业服务