遇见数据集

Supplemental data of ChEMBL-Derived Benchmark Dataset and Computational Results of Boltz-2-Based Binding Affinity Prediction

收藏
Zenodo2026-03-18 更新2026-05-26 收录
官方服务:

资源简介:

Supplemental data of a paper entitled "ChEMBL-Derived Benchmark Dataset and Computational Results of Boltz-2-Based Binding Affinity Prediction". Data S1 The amino acid sequence of the target proteins. The data contain protein targets with their UniProt IDs and target names, the corresponding PDB entries and strand IDs, and the amino acid sequences. They also contain sequence length and the sequence identity (%) for each listed target. Data S2 The curated data records and their prediction results using NIM-light. The data contain compounds associated with a target ID (UniProt ID), including each compound’s ChEMBL ID and SMILES string, along with the experimental pChEMBL value and activity type. They also contain the prediction results (confidence score, predicted affinity value and predicted binding probability) with the compound structural similarity and basic molecular properties (MW and number of RB, HA, HBD, and HBA). For each compound–target pair, the median pChEMBL values across activity types are provided as the representative values. Deduplicating compound–target pairs and their associated median pChEMBL values yields the benchmark dataset. Data S3 A KNIME workflow for creating the benchmark dataset Data S4 The calculation time of the affinity predictions. The data contain calculation times for the affinity predictions using the original Boltz-2, NIM-original, and NIM-light. The time for each target is listed with its amino acid sequence and sequence length. Data S5 Dataset A and B and the prediction results using the original Boltz-2, NIM-original, and NIM-light. The data contain experimental pChEMBL values for compound–target pairs in the dataset B. They also contain the prediction results (confidence score, predicted affinity value, and predicted binding probability) of the original Boltz-2, NIM-original, and NIM-light settings. The rows that contain prediction results of the original Boltz-2 setting correspond to the dataset A. Data S6 Benchmark dataset for predictive performance analysis and the prediction results. The data contain representative pChEMBL values (ground-truth) for compound–target pairs in the benchmark dataset for predictive performance analysis with the prediction results (confidence score, predicted affinity value, predicted binding probability) using NIM-light. For the targets used for evaluating the predictive performance without EC50 records, the representative pChEMBL values without EC50records are stored in the pChEMBL_value_median_wo_EC50 column. The data also include prediction errors, defined as the representative pChEMBL values minus the predicted affinity values, as well as compound structural similarity and basic molecular properties (MW and numbers of RB, HA, HBD, and HBA). Data S7 Figures of relationships between the pChEMBL and predicted affinity values for each target Data S8 Predictive performance metrics for each target. They were calculated using the ground truth (pChEMBL value) based on the median (e.g., Pearson_r_median) or maximum representative (e.g., Pearson_r_maximum) and on the median representative without EC50 records (e.g., Pearson_r_median_wo_EC50). The lower and upper 95% confidence bounds are included for the predictive performance metrics using the median-based ground truth. Data S9 Figures of relationships between the sequence identities and prediction errors for each target Data S10 Figures of relationships between the confidence scores and prediction errors for each target Data S11 Figures of relationships between the compound structural similarities and prediction errors for each target Data S12 Figures of relationships between the MW and prediction errors for each target Data S13 Figures of relationships between the numbers of RB and prediction errors for each target Data S14 Figures of relationships between the numbers of HA and prediction errors for each target Data S15 Figures of relationships between the numbers of HBD and prediction errors for each target Data S16 Figures of relationships between the numbers of HBA and prediction errors for each target Data S17 Prediction results for a subset of ChEMBL34. The data contain experimental pChEMBL values for compound–target pairs in a subset of ChEMBL34 with the prediction results (confidence score, predicted affinity value, predicted binding probability) using NIM-light. #File names Data S1: Data_S1_target_aaseq.csv Data S2: Data_S2_curated_data records.csv Data S3: Data_S3_KNIME_workflow_for_benchmark_dataset_preparation.knwf Data S4: Data_S4_prediction_calc_time.csv Data S5: Data_S5_dataset_AB.csv Data S6: Data_S6_benchmark_dataset_for_predictive_performance.csv Data S7: Data_S7_figures_pChEMBL_prediction_for_each_target.zip Data S8: Data_S8_predictive_performance_metrics_for_each_target.csv Data S9: Data_S9_figures_seq_identity_prediction_error_for_each_target.zip Data S10: Data_S10_figures_confidence_score_prediction_error_for_each_target.zip Data S11: Data_S11_figures_compound_sim_prediction_error_for_each_target.zip Data S12: Data_S12_figures_MW_prediction_error_for_each_target.zip Data S13: Data_S13_figures_nRB_prediction_error_for_each_target.zip Data S14: Data_S14_figures_nHA_prediction_error_for_each_target.zip Data S15: Data_S15_figures_nHBD_prediction_error_for_each_target.zip Data S16: Data_S16_figures_nHBA_prediction_error_for_each_target.zip Data S17: Data_S17_ChEMBL34_results.csv

提供机构:
Zenodo
创建时间:
2025-12-18
二维码
社区交流群
二维码
科研交流群
商业服务