遇见数据集

Data for: Connecting MHC-I-binding motifs with HLA alleles via deep learning

收藏
Mendeley Data2021-07-09 更新2026-04-09 收录
官方服务:

资源简介:

This dataset contains the research data supporting the study, “connecting MHC-I-binding motifs with HLA alleles via deep learning”. 1. MHCI_res182_seq.json: the peptide-binding cleft sequence of each MHC-I allele extracted from the IPD-IMGT/HLA database (version 3.41.0) 2. MHCI_res182_onehot.npy: the one-hot encoding of the peptide-binding cleft sequence of MHC-I allele 3. dataframe.tar.gz: this folder contains training, validation, and benchmark datasets [the common columns] 1. sequence: the peptide sequence 2. peptide_length: the length of peptide sequences 3. mhc: the MHC-I allele 4. meas: binding affinity for assay data 5. value: the value (between 0 and 1) converted from the binding affinity 6. bind: the label of binding 7. source: data source (assay, MS, random decoy, or peptide decoy) 1. train_hit.csv 1. data: measurements extracted from IEDB for the training process 2. columns 1. the common columns 2. MHCfovea: the prediction score of MHCfovea (used for ScoreCAM analysis) 2. train_decoy_{1-90}.csv 1. data: artificial decoy peptides for the training process; the data number of each file is almost equal to the number of eluted peptides 2. columns 1. the common columns 3. valid.csv 1. data: measurements extracted from IEDB and decoy peptides for validation 2. columns 1. the common columns 2. batch_size_{16, 32, 64}_and_learning_rate_{0.00100, 0.00010, 0.00001}: for optimizing hyperparameters of batch size and learning rate 3. DE_{1, 5, 10, 15, 30}_and_{30, 60, 90}: for optimizing the D-E ratio of the training and downsized dataset; the first number is the D-E ratio of the downsized dataset and the second number is the D-E ratio of the training dataset 4. benchmark.csv 1. data: eluted peptides extracted from IEDB and decoy peptides for the testing process 2. columns 1. the common columns 2. NetMHCpan4.1, MHCflurry2.0, MixMHCpred2.1, MHCfovea: the prediction score of each predictor 4. allele_expansion.tar.gz: this folder contains data for the allele expansion 1. peptides.csv: peptides used for allele expansion 2. output/{MHC-I group} 1. allele.json: alleles of the MHC-I group 2. motif.npy: binding motifs for each allele 3. prediction.npy: prediction score for each allele (the order is the same as that of peptides.csv)

本数据集包含支撑题为"基于深度学习关联主要组织相容性复合体I类(MHC-I)结合基序与人类白细胞抗原(HLA)等位基因"的研究的科研数据。 1. MHCI_res182_seq.json:从IPD-IMGT/HLA数据库(版本3.41.0)中提取的每个MHC-I等位基因的肽结合凹槽序列。 2. MHCI_res182_onehot.npy:MHC-I等位基因肽结合凹槽序列的独热编码(one-hot encoding)。 3. dataframe.tar.gz:该压缩包内含训练集、验证集与基准测试集,其通用字段如下: 1. sequence:肽序列 2. peptide_length:肽序列长度 3. mhc:MHC-I等位基因 4. meas:实验测定的结合亲和力 5. value:由结合亲和力转换得到的取值范围为0至1的数值 6. bind:结合状态标签 7. source:数据来源(包括实验测定、质谱、随机诱饵肽或诱饵肽) 该压缩包包含以下子文件: 1. train_hit.csv: 1. 数据:从IEDB(免疫表位数据库)提取的训练用测定数据 2. 字段: 1. 前述通用字段 2. MHCfovea:MHCfovea的预测得分(用于ScoreCAM分析) 2. train_decoy_{1-90}.csv: 1. 数据:用于训练流程的人工合成诱饵肽,各文件的数据量与洗脱肽的数量基本相当 2. 字段: 1. 前述通用字段 3. valid.csv: 1. 数据:用于验证流程、从IEDB提取的测定数据与诱饵肽 2. 字段: 1. 前述通用字段 2. batch_size_{16, 32, 64}_and_learning_rate_{0.00100, 0.00010, 0.00001}:用于优化批量大小与学习率的超参数 3. DE_{1, 5, 10, 15, 30}_and_{30, 60, 90}:用于优化训练集与缩小规模数据集的D-E比例;其中第一个数值为缩小规模数据集的D-E比例,第二个数值为训练集的D-E比例 4. benchmark.csv: 1. 数据:用于测试流程、从IEDB提取的洗脱肽与诱饵肽 2. 字段: 1. 前述通用字段 2. NetMHCpan4.1、MHCflurry2.0、MixMHCpred2.1、MHCfovea:各预测工具的预测得分 4. allele_expansion.tar.gz:该压缩包包含等位基因扩增相关数据 1. peptides.csv:用于等位基因扩增的肽序列 2. output/{MHC-I组}: 1. allele.json:该MHC-I组的等位基因信息 2. motif.npy:各等位基因的结合基序数据 3. prediction.npy:各等位基因的预测得分(顺序与peptides.csv中的肽序列一致)

创建时间:
2021-07-09
二维码
社区交流群
二维码
科研交流群
商业服务