遇见数据集

Enhancing the resolvability of cryo-EM maps in protein-ligand complexes using deep learning

收藏
Zenodo2026-08-04 更新2026-08-13 收录
官方服务:

资源简介:

Overview This repository contains the dataset, metadata, data-splitting manifests, and source training code for the machine learning framework introduced in the manuscript: "Enhancing the resolvability of cryo-EM maps in protein-ligand complexes using deep learning" Training Dataset & Metadata ml_dataset_FINAL_BALANCED.h5: It contains 64×64×64 subgrid crops at a pixel size of 0.5 A˚ spanning experimental density maps, simulated target maps. To prevent over-representation of common cofactors (e.g., ATP), instances are capped at a maximum of 50 samples per unique SMILES string. FINAL_dataset_inventory_BALANCED.xlsx: A metadata for every instance in the H5 file. It lists parent PDB IDs, target ligand 3-letter codes, canonical SMILES strings, and structural classification labels. Data Split exact_ligand_clusters_FINAL.csv: Contains cluster indices grouping data points based on protein sequence similarity. This file ensures that homologous or identical proteins are clustered together. dataset_split_manifest.csv: The finalized Train, Validation, and Test split allocation sheet generated via the sequence similarity clusters, ensuring a unbiased benchmark for model performance. Deep Learning Code Base 04_train_with_datasplit.py: The primary executable training script. architecture.py: The Python script containing the neural network topology. utils_common.py: Shared helper utility code containing pipeline functions. Visualizations & Reproducibility ChimeraX Sessions (Figure 5 Data Points): Pre-saved 3D visualization workspace environments. These sessions map directly to the structural results showcased in Figure 5 of the accompanying manuscript.

提供机构:
Zenodo
创建时间:
2026-08-04
二维码
社区交流群
二维码
科研交流群
商业服务