Additional file 1: of Integrative analysis and machine learning on cancer genomics data using the Cancer Systems Biology Database (CancerSysDB)
收藏资源简介:
The source code of the database queries and workflow scripts for the three use cases reported in the paper. The results can be reproduced using the query results and analysis scripts provided. File query1.csv contains the barcodes of all samples for which mutation data do exist. File query2.csv contains the barcodes of all samples which carry a mutation in the gene of interest. Finally, query3.csv contains the survival data (according to Fig. 1a), a list of all mutations of patients in the cohort of interest (according to Fig. 1b), or a list of all genomic segments with aberrant copy number in the cohort of interest (according to Fig. 1c). There are small discrepancies between the number of patients with mutation data and the number of patients with survival data (Fig. 1a) and copy number data (Fig. 1c). (ZIP 4981 kb)
本数据集包含论文中报告的三个用例对应的数据库查询与工作流脚本的源代码。依托所提供的查询结果与分析脚本,可复现相关实验结果。其中query1.csv存储了所有存在突变数据的样本条形码(barcodes);query2.csv存储了所有在目标基因中携带突变的样本条形码;而query3.csv则包含三类数据:一是对应图1a的生存数据,二是对应图1b的目标队列中所有患者的突变列表,三是对应图1c的目标队列中所有拷贝数异常的基因组区段列表。存在突变数据的患者数量与带有生存数据(图1a)及拷贝数数据(图1c)的患者数量之间存在小幅差异。(压缩包大小:4981 KB)



