Breast cancer dataset used in: Biologically informed NeuralODEs for genome-wide regulatory dynamics
收藏资源简介:
The original data set comes from a cross-sectional breast cancer study (GEO accession GSE7390) consisting of microarray expression values for 22000 genes from 198 breast cancer patients, that we sorted along a pseudotime axis. We noted that the same data set was also used in the PROB paper (Sun, X., et al 2021, Inferring latent temporal progression and regulatory networks from cross-sectional transcriptomic data of cancer samples). PROB is a GRN inference method that infers a random-walk-based pseudotime to sort cross-sectional samples and reconstruct the GRN. For consistency and convenience in pseudotime inference, we obtained the same version of this data that was already preprocessed and sorted by PROB. This was shared with us by the authors of PROB. We have uploaded the shared files here, as well as the versions obtained after pre-processing to apply PHOENIX.
本原始数据集源自一项横断面乳腺癌研究,其对应基因表达综合数据库(Gene Expression Omnibus,GEO)登录号为GSE7390,包含198名乳腺癌患者的22000个基因的微阵列表达值,我们已将该数据集沿伪时间(pseudotime)轴完成排序。我们注意到该数据集也曾被用于PROB相关研究(Sun X等,2021年,《从癌症样本的横断面转录组数据中推断潜在时间进程与调控网络》)。PROB是一种基因调控网络(Gene Regulatory Network,GRN)推断方法,其通过基于随机游走的伪时间分析对横断面样本进行排序,并重构基因调控网络。为保证伪时间推断的一致性与便利性,我们获取了该数据集经PROB预处理并排序后的同一版本,该版本由PROB方法的原作者共享给我们。我们已将该共享文件以及经预处理以适配PHOENIX的数据集版本上传至此处。



