simulated data from:MILP: a meta learning-based interpretable deep learning framework for multi-omics genomic prediction
收藏资源简介:
This simulation dataset was used for MILP: a meta Learning-Based interpretable deep learning framework for multi-omics genomic prediction The dataset encompasses simulated data pertaining to six distinct scenarios. These six simulated scenarios are elaborated as follows: scenarios characterized by heritabilities of 0.1 (h2=0.1) and 0.3 (h2=0.3), along with the presence of 100 (QTL100), 150 (QTL150), and 200 (QTL200) Quantitative Trait Loci (QTLs). The six sets of simulated data are respectively housed in two folders, namely simulated_data1-3-5 and simulated_data2-4-6. The folder simulated_data1-3-5 contains three files, corresponding to the simulated data for a heritability of 0.1 and QTL numbers of 100, 150, and 200 respectively. In contrast, the folder simulated_data2-4-6 contains the simulated data for a heritability of 0.3 and QTL numbers of 100, 150, and 200. The sequential numbering is presented as follows: QTLs h2 0.1 0.3 100 simulated_data_1 simulated_data_2 150 simulated_data_3 simulated_data_4 200 simulated_data_5 simulated_data_6 Each simulated data file contains genotype data and phenotype data. For more information about the simulated data, please refer to the research of "MILP: a meta learning-based interpretable deep learning framework for multi-omics genomic prediction". In the content format of the file, the first line is the ID number of the individual, the second line is the corresponding phenotypic value of the individual. Starting from the third line, it is the genotype data, which is stored in the format of 0, 1, and 2. Among them, "0" represents the homozygous reference genotype, that is, both alleles are the reference sequence; "1" represents the heterozygous genotype, which contains one reference allele and one variant allele; "2" corresponds to the homozygous variant genotype, where both alleles are in the variant form. All six simulated data files are stored in the above format. The QMSim folder contains parameter card files corresponding to the respective file names. The corresponding output files can be generated by executing these parameter cards in the QMSim software.
本仿真数据集用于MILP:一种面向多组学基因组预测的基于元学习的可解释深度学习框架(meta Learning-Based interpretable deep learning framework for multi-omics genomic prediction)。该数据集涵盖6种不同场景下的仿真数据,具体场景如下:分别为遗传力0.1(h²=0.1)、0.3(h²=0.3),以及包含100个(QTL100)、150个(QTL150)、200个(QTL200)数量性状基因座(Quantitative Trait Loci,QTL)的场景。 6组仿真数据分别存放于两个文件夹中,即simulated_data1-3-5与simulated_data2-4-6。其中文件夹simulated_data1-3-5包含3个文件,分别对应遗传力0.1、QTL数量为100、150、200的仿真数据;与之对应,文件夹simulated_data2-4-6则包含遗传力0.3、QTL数量为100、150、200的仿真数据,对应关系如下: QTL数量为100时,对应simulated_data_1(遗传力0.1)与simulated_data_2(遗传力0.3);QTL数量为150时,对应simulated_data_3(遗传力0.1)与simulated_data_4(遗传力0.3);QTL数量为200时,对应simulated_data_5(遗传力0.1)与simulated_data_6(遗传力0.3)。 每份仿真数据文件均包含基因型数据与表型数据。如需了解更多仿真数据相关细节,请参阅论文《MILP: a meta learning-based interpretable deep learning framework for multi-omics genomic prediction》。 文件内容格式如下:首行为个体编号,次行对应个体的表型值;自第三行起为基因型数据,以0、1、2三种数值存储。其中,“0”代表纯合参考基因型,即两个等位基因均为参考序列;“1”代表杂合基因型,包含一个参考等位基因与一个变异等位基因;“2”对应纯合变异基因型,即两个等位基因均为变异形式。 全部6份仿真数据文件均采用上述格式存储。 QMSim文件夹中包含与各文件名对应的参数卡文件,通过在QMSim软件中运行这些参数卡即可生成对应的输出文件。



