ACT representative primary microarray samples
收藏资源简介:
Arabidopsis thaliana ATH1 raw microarray data (CEL files) were download from ArrayExpress, GEO and NASCArrays. After duplicate and corrupt sample removal, using an in-house PHP script, our dataset consisted of 19,887 unique microarray samples from 1390 studies. Finally, after quality control, 6933 distinct, wild-type, healthy samples were selected. Pairwise sample correlations were calculated using Pearson Correlation Coefficient, using the expression values of 21273 non-obsolete Arabidopsis thaliana genes, in the 6933 previously selected samples. A sample distance matrix was created using the d = 1 – r formula and a sample correlation tree of 6933 sample-leaves was created in Newick format, based on the distance matrix. Then, the tree of 6933 sample-leaves, was programmatically pruned in an iterative procedure using an in-house algorithm trimming close leaves, where in each iteration the leaf with the shortest distance to its first common node was trimmed, leaving 3500 leaves which represent the most distinct samples.
拟南芥(Arabidopsis thaliana)ATH1芯片原始微阵列数据(CEL文件)下载自ArrayExpress、GEO与NASCArrays三大数据库。首先通过自研PHP脚本完成重复样本与损坏样本的剔除工作,最终得到包含1390项研究、共计19887个唯一微阵列样本的数据集。随后经过质量控制流程,筛选出6933个独立的野生型健康样本。针对上述6933个选定样本,基于21273个未废弃拟南芥基因的表达值,采用皮尔逊相关系数(Pearson Correlation Coefficient)计算样本间的两两相关性。通过公式d = 1 – r构建样本距离矩阵,并基于该矩阵生成包含6933个样本叶节点的Newick格式样本相关聚类树。随后通过自研算法对该聚类树执行迭代剪枝流程:每次迭代中移除与最近共同节点距离最短的叶节点,最终保留3500个最具代表性的独特样本,对应3500个叶节点。



