Expression Data recompute of selected GEO-deposited RNA-Seq data of HMEC-1 cell lines
收藏资源简介:
We aligned and quantified RNA-Seq data present in GEO regarding HMEC-1 cell lines with a standardized pipeline to homogenize data preprocessing for downstream applications. All uploaded files are UTF-8, .csv-formatted matrices. The *_expected_count.csv.gz files are unlogged, raw expression counts as reported by rsem-quantify-expression with the 'expected counts' feature. The associated *_metadata.csv.gz files contain metadata pertinent to each column of the corresponding expression matrix.Some metadata files may have more rows than the associated number of columns. This is for series that were only partially RNA-Seq based (e.g. combinated RNA-Seq plus miRNA-Seq samples in the same GEO accession ID). Metadata columns are derived from GEO series files, and follow their definitions. See each GEO entry directly to determine metadata meaning. Each recompute has at least the gene_id column holding Ensembl Gene IDs. The remaining columns are ENA run accession IDs of the specific recomputed samples.Each associated metadata has at least the following columns: geo_sample: The GEO sample ID of the sample. geo_series: The GEO series ID of the sample. ena_sample: The ENA sample ID of the sample. ena_run: The ENA run accession ID of the sample, to be cross-referenced with the expression matrices. The remaining columns are derived from GEO metadata files and other ENA-provided data. Please refer to the x.FASTQ package for more information (https://github.com/TCP-Lab/x.FASTQ).Reference genome was downloaded from Ensembl, version hg38. STAR was used to create the index genome with overhang set to 149.The different datasets where generated over a long period of time trough a variety of different versions of x.FASTQ. However, the versions of the softwares that acted on the files themselves (e.g. STAR, rsem, etc...) were unchanged. Changelog Version 2: Added GSE139947 and replaced faulty GSE244042 and GSE199978, which were missing some samples.
本研究采用标准化分析流程,对基因表达综合数据库(Gene Expression Omnibus,GEO)中HMEC-1细胞系的RNA测序(RNA-Seq)数据进行序列比对与定量,以统一数据预处理标准,适配各类下游分析场景。 所有上传文件均采用UTF-8编码,格式为逗号分隔值(CSV)矩阵文件。其中*_expected_count.csv.gz文件为未经过对数转换的原始表达计数数据,由rsem-quantify-expression工具的"expected counts"功能生成。与之配套的*_metadata.csv.gz文件则包含对应表达矩阵各列的相关元数据。 部分元数据文件的行数可能多于对应表达矩阵的列数,此类情况多见于仅部分样本采用RNA-Seq技术的数据集(例如同一GEO登录号下同时包含RNA-Seq与microRNA测序(miRNA-Seq)样本的数据集)。 元数据列均源自GEO系列文件,并遵循其定义规范,可直接查阅对应GEO条目以明确元数据的具体含义。 每个重计算得到的表达矩阵至少包含gene_id列,该列存储Ensembl基因ID(Ensembl Gene IDs)。其余列则为对应重计算样本的欧洲核苷酸档案馆(European Nucleotide Archive,ENA)测序运行登录号。 配套的元数据文件至少包含以下列: geo_sample:对应样本的GEO样本登录号。 geo_series:对应样本的GEO系列登录号。 ena_sample:对应样本的ENA样本登录号。 ena_run:对应样本的ENA测序运行登录号,可与表达矩阵进行交叉比对。 其余列均源自GEO元数据文件及欧洲核苷酸档案馆提供的其他数据。如需了解更多信息,请参阅x.FASTQ工具包(https://github.com/TCP-Lab/x.FASTQ)。 参考基因组下载自Ensembl数据库,版本为hg38。采用STAR工具构建基因组索引,设置overhang参数为149。本数据集的不同子数据集在较长周期内通过多个版本的x.FASTQ工具生成,但用于处理原始文件的工具(如STAR、rsem等)版本始终保持一致。 更新日志 版本2:新增GSE139947数据集,并替换存在样本缺失问题的GSE244042与GSE199978数据集。



