Meta analysis of public drought gene expression data in plants
收藏资源简介:
Physiologically relevant drought stress is difficult to apply consistently, and the heterogeneity in experimental design, growth conditions, and sampling schemes make it challenging to compare water deficit studies in plants. Here, we re-analyzed hundreds of drought gene expression experiments across diverse model and crop species and quantified the variability across studies. We found that drought studies are surprisingly uncomparable, even when accounting for differences in genotype, environment, drought severity, and method of drying. Many studies, including most Arabidopsis work, lack high-quality phenotypic and physiological datasets to accompany gene expression, making it impossible to assess the severity or in some cases the occurrence of water deficit stress events. From these datasets, we developed supervised learning classifiers that can accurately predict if RNA-seq samples have experienced a physiologically relevant drought stress, and suggest this can be used as a quality c..., We assembled a database of drought RNAseq data in Arabidopsis, soybean, tomato, rice, maize, and rice from the NCBI sequence read archive (SRA). Bulk data was retrieved using a series of drought or heat stress related keywords with the SRA Advanced Search Builder. The following metadata was collected for each experiment: tissue type(s), developmental stage, environment (e.g, greenhouse, field, growth chamber etc), media type, duration of stress, mechanism of drying, associated physiology datasets, genotype, number of timepoints, and number of replicates. 112 studies had a linked publication in the NCBI metadata and 130 had no associated publication across all 6 species. Similar metadata was retrieved for individual SRA samples along with a binary classification of treatment (drought or control) where possible. Metadata was retrieved from the SRA and associated publications, but the lack of publications and ambiguity in some labels led to a high degree of missing or sparse metadata..., , # Meta analysis of public drought gene expression data in plants [https://doi.org/10.5061/dryad.7sqv9s50g](https://doi.org/10.5061/dryad.7sqv9s50g) Here, we re-analyzed hundreds of drought gene expression experiments across diverse model and crop species and quantified the variability across studies. We assembled a database of drought RNAseq data in Arabidopsis, soybean, tomato, rice, maize, and rice from the NCBI sequence read archive (SRA). Bulk data was retrieved using a series of drought stress related keywords with the SRA Advanced Search Builder. The following metadata was collected for each experiment: tissue type(s), developmental stage, environment (e.g, greenhouse, field, growth chamber etc), media type, duration of stress, mechanism of drying, associated physiology datasets, genotype, number of timepoints, and number of replicates. 112 studies had a linked publication in the NCBI metadata and 130 had no associated publication across all 6 species. Similar metadata was retr..., ,
生理相关性干旱胁迫难以实现标准化施加,而实验设计、生长条件与采样方案存在的异质性,使得植物水分亏缺相关研究间的对比极具挑战性。 本研究对涵盖多种模式植物与作物物种的数百项干旱基因表达实验进行了重新分析,并量化了不同研究间的变异程度。结果令人意外地发现,即便考虑基因型、生长环境、干旱强度以及干旱诱导方式的差异,现有干旱相关研究仍难以相互比较。包括多数拟南芥(Arabidopsis)研究在内的诸多项目,均缺乏与基因表达数据配套的高质量表型与生理数据集,导致无法评估水分亏缺胁迫的强度,部分情况下甚至无法确认胁迫是否发生。基于上述数据集,本研究开发了监督学习分类器,可精准预测RNA测序(RNA-seq)样本是否经历过生理相关性干旱胁迫,并提出该分类器可作为质量控制[原文截断]。 本研究从美国国家生物技术信息中心(National Center for Biotechnology Information, NCBI)的序列读取归档库(Sequence Read Archive, SRA)中,收集了拟南芥、大豆、番茄、水稻、玉米与水稻的干旱RNA-seq数据,构建了对应的数据库。通过SRA高级搜索构建器,使用一系列干旱或热胁迫相关关键词批量获取了数据。每项实验均收集了如下元数据:组织类型、发育阶段、生长环境(如温室、大田、生长箱等)、培养基类型、胁迫持续时长、干旱诱导机制、配套生理数据集、基因型、时间点数量以及生物学重复次数。在所有6个物种中,有112项研究在NCBI元数据中关联了已发表文献,另有130项研究未关联任何公开出版物。针对单个SRA样本,我们也收集了类似的元数据,并在可行的情况下为其标注了处理组(干旱胁迫或对照组)的二分类标签。元数据从SRA数据库及关联文献中获取,但由于部分研究未发表、部分标注存在歧义,导致元数据存在大量缺失或稀疏的情况[原文截断]。 # 植物公共干旱基因表达数据的元分析 [https://doi.org/10.5061/dryad.7sqv9s50g](https://doi.org/10.5061/dryad.7sqv9s50g) 本研究对涵盖多种模式植物与作物物种的数百项干旱基因表达实验进行了重新分析,并量化了不同研究间的变异程度。本研究从美国国家生物技术信息中心(National Center for Biotechnology Information, NCBI)的序列读取归档库(Sequence Read Archive, SRA)中,收集了拟南芥、大豆、番茄、水稻、玉米与水稻的干旱RNA-seq数据,构建了对应的数据库。通过SRA高级搜索构建器,使用一系列干旱胁迫相关关键词批量获取了数据。每项实验均收集了如下元数据:组织类型、发育阶段、生长环境(如温室、大田、生长箱等)、培养基类型、胁迫持续时长、干旱诱导机制、配套生理数据集、基因型、时间点数量以及生物学重复次数。在所有6个物种中,有112项研究在NCBI元数据中关联了已发表文献,另有130项研究未关联任何公开出版物。类似的元数据获取工作也针对单个SRA样本开展,[原文截断]。



