遇见数据集

KinFragLib_PocketEnum: Novel Kinase Ligand Generation Using Subpocket-Based Docking

收藏
Zenodo2025-12-17 更新2026-05-26 收录
官方服务:

资源简介:

Project description. Protein kinases are crucial for signal transduction by catalyzing phosphorylation. Dysregulation is linked to diseases like cancer, autoimmune disorders, and Alzheimer's, making them key drug targets. High interest in this protein family has yielded large amounts of available structure and ligand data. Fragment-based drug design has shown promise in developing kinase inhibitors but often ignores kinase-specific knowledge.The KinFragLib library addresses this by fragmenting co-crystallized kinase ligands based on functionally relevant subpockets, resulting in a library of 7486 fragments from 2553 kinase-ligand complexes. Despite the extensive chemical space spanned by this library, enumerating all possible recombinations is computationally infeasible. To tackle this, we developed an automated Python pipeline that orchestrates fragment growing within the binding site of any protein kinase of interest, employing SeeSAR's docking engine. A guided template docking search, informed by subpocket-specific information, reduces the combinatorial space, ensuring efficient ligand generation. The pipeline is available at https://github.com/volkamerlab/KinFragLib_PocketEnum. In a case study, we applied our subpocket-based docking approach on a PKA kinase, generating 33,758 compounds. 1154 of these compounds were considered drug-like according to the Rule of Five and have HYDE affinity scores in the nM range. When compared to the space covered by ChEMBL33, these compounds exhibit novel chemicalmatter of more than 99.1% considering a similarity cut-off greater than or equal to 0.9. Case study: Subpocket-based docking to identify potential PKA inhibitors. The dataset was generated as a case study of the subpocket-based docking approach (https://github.com/volkamerlab/KinFragLib_PocketEnum) against the hamster PKA (PDB code: 5n1f) using the kinase-focused library KinFragLib (v2.1.0). Note that BioSolveIT's SeeSAR (3D desktop modeling platform to prepare the protein input files) and the following SeeSAR command-line tools were used here: FlexX - for docking HYDE - for scoring and optimization. 1. Input data config/5n1f/ settings.json: program configuration AP.flexx, FP.flexx, GA.flexx: subpocket-specific docking (FlexX) configuration files AP.hydescorer, FP.hydescorer, GA.hydescorer: subpocket-specific scoring & optimization (HYDE) configuration files Note: for more information, refer to the two getting started notebooks in https://github.com/volkamerlab/KinFragLib_PocketEnum/notebooks. 2. Generated data The resulting data from running the subpocket-based docking pipeline on the given settings (config/5n1f/) results_5n1f_25_02/ 5n1f_out.json: program statistics, such as the number of generated docking poses per subpocket iteration 5n1f/ results.sdf: generated ligands, comprising the best-scored docking pose of each recombined ligand that covers either only subpockets AP and FP, or all subpockets (AP, FP, and GA) SP0.sdf, SP1.sdf, and SP2.sdf: all generated docking poses of all ligands per subpocket iteration violations_SP0.sdf, violations_SP1.sdf and violations_SP2.sdf: docking poses of ligands that were disregarded since the change in their pose exceeded the given (RMSD) threshold after HYDE-optimization proposed_ligands.sdf: the seven selected ligands for synthesis from results.sdf adapted_mols.sdf: modified versions of the ligands from proposed_ligands.sdf for synthesis results_chembl.csv: (RDKit's topological) fingerprint-based similarity comparison against ChEMBL33 based on the Tanimoto coefficient (output of the comparison script chembl_database_comp.py) kinase_compare.csv: (RDKit's topological) fingerprint-based similarity comparison against the cocrystallized kinase ligands of the PDB based on the Tanimoto coefficient (output of the comparison script pdb_db_comp.py) pka_compare.csv: (RDKit's topological) fingerprint-based similarity comparison against the cocrystallized PKA ligands (of arbitrary organisms) of the PDB based on the Tanimoto coefficient (output of the comparison script pdb_db_comp.py) hamster_pka_compare.csv: (RDKit's topological) fingerprint-based similarity comparison against the cocrystallized hamster PKA ligands of the PDB based on the Tanimoto coefficient (output of the comparison script pdb_db_comp.py) pka_compare_synth.csv: (RDKit's topological) fingerprint-based similarity comparison of our compounds that were selected for synthesis and their modified versions against the cocrystallized PKA ligands (of arbitrary organisms) of the PDB based on the Tanimoto coefficient (output of the comparison script pdb_db_comp.py) hamster_pka_compare_synth.csv: (RDKit's topological) fingerprint-based similarity comparison of our compounds that were selected for synthesis and their modified versions against the cocrystallized hamster PKA ligands (of arbitrary organisms) of the PDB based on the Tanimoto coefficient (output of the comparison script pdb_db_comp.py) Usage The configuration data can be used to reproduce the generated dataset, while the dataset can be used to run the notebooks available on https://github.com/volkamerlab/KinFragLib_PocketEnum. Clone the KinFragLib_PocketEnum repository. Download the two tar files provided here. Extract the archive content into your local KinFragLib_PocketEnum folder and run the pipeline or the data analysis notebooks. Configuration files: tar -xvf config.tar -C /path_to_kinfraglib_pocket_enum/config Generated data: tar -xvf results_5n1f_25_02.tar -C /path_to_kinfraglib_pocket_enum Citation. This dataset is part of the subpocket-based docking publication: Buchthal K, Kramer PL, Hubach D, Bach G, Wagner N, Krieger J, et al. Novel Kinase Ligand Generation Using Subpocket-Based Docking. ChemRxiv. 2025; doi:10.26434/chemrxiv-2025-f9ctg. This content is a preprint and has not been peer-reviewed.

提供机构:
Zenodo
创建时间:
2025-12-17
二维码
社区交流群
二维码
科研交流群
商业服务