遇见数据集

GEO gene expression dataset recompute for selected tumor samples

收藏
Zenodo2024-05-15 更新2026-05-26 收录
官方服务:

资源简介:

We aligned and quantified RNA-Seq data present in GEO with a standardized pipeline to homogenize data preprocessing for downstream applications. All uploaded files are UTF-8, .csv-formatted matrices. The *_expected_count.csv.gz files are unlogged, raw expression counts as reported by rsem-quantify-expression (see details below). The associated *_metadata.csv.gz files contain metadata pertinent to each column of the corresponding expression matrix.Some metadata files may have more rows than the associated number of columns. This is for series that were only partially RNA-Seq based (e.g. combinated RNA-Seq plus miRNA-Seq samples in the same GEO accession ID). Metadata columns are derived from GEO series files, and follow their definitions. See each GEO entry directly to determine metadata meaning. Each recompute has at least the gene_id column holding Ensembl Gene IDs. The remaining columns are ENA run accession IDs of the specific recomputed samples.Each associated metadata has at least the following columns: geo_accession: The GEO sample ID of the sample. ena_sample: The ENA sample ID of the sample. ena_run: The ENA run accession ID of the sample, to be cross-referenced with the expression matrices. The remaining columns are derived from GEO metadata files and other ENA-provided data. Please refer to the x.FASTQ package for more information. Pipeline Details The alignment and quantification was made with the x.FASTQ tool available on Github installed locally on an Arch Linux machine on commit 3a93dd77a70df59c74f7b15216c26f12cd918e81 running the Linux 6.7.8-zen1-1-zen kernel with a 11th Gen Intel i7-1185G7 (8) CPU and a Intel TigerLake-LP GT2 [Iris Xe Graphics] GPU. Please note that no sample filtering or omissions were done based on sample quality or sequencing depth. However, sensible trimming (e.g. low-quality bases and common adapters) was performed on all the samples. Reference genome was downloaded from Ensembl, version hg38. STAR was used to create the index genome with overhang set to 149.

我们采用标准化分析流程,对基因表达汇编(Gene Expression Omnibus, GEO)中的现有RNA-Seq数据进行比对与定量,以统一数据预处理规范,适配各类下游分析需求。 所有上传文件均为UTF-8编码的.csv格式矩阵文件。其中*_expected_count.csv.gz文件为未经对数转换的原始表达计数数据,由rsem-quantify-expression工具生成(详见下文说明)。与之配套的*_metadata.csv.gz文件包含对应表达矩阵各列的相关元数据。部分元数据文件的行数多于对应表达矩阵的列数,这是因为部分数据集仅部分样本采用RNA-Seq测序(例如同一GEO收录编号下同时包含RNA-Seq与miRNA-Seq样本的组合数据集)。 元数据列均源自GEO数据集文件,遵循其定义规范。如需了解元数据的具体含义,请直接查阅对应GEO条目。 每份重处理后的表达矩阵至少包含gene_id列,其中存储Ensembl(Ensembl)基因ID;其余列均为对应重处理样本的欧洲核苷酸档案馆(European Nucleotide Archive, ENA)运行收录编号。每份配套的元数据文件至少包含以下列: geo_accession:样本的GEO样本ID。 ena_sample:样本的ENA样本ID。 ena_run:样本的ENA运行收录编号,可与表达矩阵进行交叉比对。 其余列均源自GEO元数据文件及其他ENA提供的数据。如需了解更多信息,请参考x.FASTQ工具包。 ## 流程详情 本次比对与定量分析使用部署于Arch Linux主机的本地GitHub开源工具x.FASTQ完成,对应提交版本为commit 3a93dd77a70df59c74f7b15216c26f12cd918e81,运行内核为Linux 6.7.8-zen1-1-zen,CPU为第11代Intel i7-1185G7(8核),GPU为Intel TigerLake-LP GT2 [Iris Xe Graphics]。请注意,本次分析未基于样本质量或测序深度对样本进行过滤或剔除,但已对所有样本执行了合理的序列修剪操作(例如去除低质量碱基与通用接头序列)。 参考基因组下载自Ensembl数据库(Ensembl),版本为hg38。使用STAR(STAR)工具构建基因组索引,设置overhang参数为149。

提供机构:
Zenodo
创建时间:
2024-05-13
二维码
社区交流群
二维码
科研交流群
商业服务