Data from: Machine learning-based differential network analysis: a study of stress-responsive transcriptomes in Arabidopsis thaliana
收藏资源简介:
Machine learning (ML) is an intelligent data mining technique that builds a prediction model based on the learning of prior knowledge to recognize patterns in large-scale data sets. We present an ML-based methodology for transcriptome analysis via comparison of gene coexpression networks, implemented as an R package called machine learning–based differential network analysis (mlDNA) and apply this method to reanalyze a set of abiotic stress expression data in Arabidopsis thaliana. The mlDNA first used a ML-based filtering process to remove nonexpressed, constitutively expressed, or non-stress-responsive “noninformative” genes prior to network construction, through learning the patterns of 32 expression characteristics of known stress-related genes. The retained “informative” genes were subsequently analyzed by ML-based network comparison to predict candidate stress-related genes showing expression and network differences between control and stress networks, based on 33 network topological characteristics. Comparative evaluation of the network-centric and gene-centric analytic methods showed that mlDNA substantially outperformed traditional statistical testing–based differential expression analysis at identifying stress-related genes, with markedly improved prediction accuracy. To experimentally validate the mlDNA predictions, we selected 89 candidates out of the 1784 predicted salt stress–related genes with available SALK T-DNA mutagenesis lines for phenotypic screening and identified two previously unreported genes, mutants of which showed salt-sensitive phenotypes.
机器学习(Machine Learning, ML)是一种智能数据挖掘技术,通过学习先验知识构建预测模型以识别大规模数据集内的模式。本研究提出了一种基于机器学习的转录组分析方法,该方法通过比较基因共表达网络(gene coexpression networks)实现,并被封装为一款名为基于机器学习的差异网络分析(machine learning–based differential network analysis, mlDNA)的R软件包;随后将该方法应用于重新分析一组拟南芥(Arabidopsis thaliana)的非生物胁迫表达数据集。mlDNA首先采用基于机器学习的筛选流程,在构建网络前先剔除未表达、组成型表达(constitutively expressed)或无胁迫响应的“非信息性”基因,该流程通过学习已知胁迫相关基因的32种表达特征模式完成。后续,基于33种网络拓扑特征(network topological characteristics),研究人员通过基于机器学习的网络比较方法对保留的“信息性”基因进行分析,以预测在对照组与胁迫组网络中存在表达与网络差异的候选胁迫相关基因。通过对网络中心型(network-centric)与基因中心型(gene-centric)分析方法的对比评估,结果显示mlDNA在识别胁迫相关基因方面显著优于传统基于统计检验的差异表达分析方法,预测精度得到大幅提升。为实验验证mlDNA的预测结果,我们从1784个预测的盐胁迫相关基因中,挑选出89个拥有可用SALK T-DNA突变体库的候选基因开展表型筛选,最终鉴定出两个此前未被报道的基因,其突变体呈现盐敏感表型。



