遇见数据集

Expression Data recompute of selected GEO-deposited RNA-Seq data of HMEC-1 cell lines

收藏
Zenodo2025-02-11 更新2026-05-26 收录
官方服务:

资源简介:

We aligned and quantified RNA-Seq data present in GEO regarding HMEC-1 cell lines with a standardized pipeline to homogenize data preprocessing for downstream applications. All uploaded files are UTF-8, .csv-formatted matrices. The *_expected_count.csv.gz files are unlogged, raw expression counts as reported by rsem-quantify-expression with the 'expected counts' feature. The associated *_metadata.csv.gz files contain metadata pertinent to each column of the corresponding expression matrix.Some metadata files may have more rows than the associated number of columns. This is for series that were only partially RNA-Seq based (e.g. combinated RNA-Seq plus miRNA-Seq samples in the same GEO accession ID). Metadata columns are derived from GEO series files, and follow their definitions. See each GEO entry directly to determine metadata meaning. Each recompute has at least the gene_id column holding Ensembl Gene IDs. The remaining columns are ENA run accession IDs of the specific recomputed samples.Each associated metadata has at least the following columns: geo_sample: The GEO sample ID of the sample. geo_series: The GEO series ID of the sample. ena_sample: The ENA sample ID of the sample. ena_run: The ENA run accession ID of the sample, to be cross-referenced with the expression matrices. The remaining columns are derived from GEO metadata files and other ENA-provided data. Please refer to the x.FASTQ package for more information (https://github.com/TCP-Lab/x.FASTQ).Reference genome was downloaded from Ensembl, version hg38. STAR was used to create the index genome with overhang set to 149.The different datasets where generated over a long period of time trough a variety of different versions of x.FASTQ. However, the versions of the softwares that acted on the files themselves (e.g. STAR, rsem, etc...) were unchanged, and reported below:

本数据集采用标准化流程对基因表达综合数据库(Gene Expression Omnibus,GEO)中收录的HMEC-1细胞系RNA测序(RNA-Seq)数据进行序列比对与定量分析,以统一数据预处理流程,适配下游分析需求。 所有上传文件均为UTF-8编码的逗号分隔值(CSV)格式矩阵文件。其中带`*_expected_count.csv.gz`后缀的文件为未经过对数转换的原始表达计数结果,由`rsem-quantify-expression`工具的“expected counts”功能生成。对应的`*_metadata.csv.gz`元数据文件包含与对应表达矩阵各列相关的元数据信息。 部分元数据文件的行数会多于对应表达矩阵的列数,这是因为部分GEO系列数据集仅部分样本采用了RNA测序(RNA-Seq)技术(例如同一GEO登录号下同时包含RNA测序与小RNA测序(miRNA-Seq)样本的情况)。 元数据列均源自GEO系列数据集文件,并遵循其原始定义。如需了解元数据的具体含义,请直接查阅对应GEO条目。 每个重计算得到的表达矩阵至少包含`gene_id`列,存储Ensembl基因ID(Ensembl Gene IDs)。其余列均为对应重计算样本的欧洲核苷酸档案库(European Nucleotide Archive,ENA)运行登录号。每个配套元数据文件至少包含以下列: `geo_sample`:对应样本的GEO样本编号 `geo_series`:对应样本的GEO系列编号 `ena_sample`:对应样本的欧洲核苷酸档案库(ENA)样本编号 `ena_run`:对应样本的欧洲核苷酸档案库(ENA)运行登录号,可与表达矩阵进行交叉比对 其余元数据列均源自GEO元数据文件及欧洲核苷酸档案库提供的其他数据。如需了解更多信息,请查阅`x.FASTQ`工具包(https://github.com/TCP-Lab/x.FASTQ)。 参考基因组下载自Ensembl数据库(Ensembl),版本为hg38。我们使用STAR工具构建基因组索引,并将overhang参数设置为149。 本数据集的不同子数据集是在较长时间内通过多个版本的`x.FASTQ`工具生成的,但用于处理原始文件的工具(如STAR、rsem等)版本均保持固定,各工具版本信息如下:

提供机构:
Zenodo
创建时间:
2025-02-03
二维码
社区交流群
二维码
科研交流群
商业服务