Berberine–Ulcerative Colitis Drug–Target Interaction Dataset
收藏资源简介:
BU-DTI-BerbUC565K is a large-scale drug–target interaction dataset designed to support graph learning, network pharmacology analysis, and docking-based prediction of berberine targets relevant to ulcerative colitis. The dataset contains 565,605 compound–protein pairs spanning 3,200 candidate protein targets, where each sample represents a biologically plausible interaction scenario between berberine and a potential disease-associated target. Each record integrates multi-domain features grouped into five categories: 1. Compound Chemical Properties:Physicochemical descriptors including molecular weight, logP, polar surface area (TPSA), hydrogen bond donors/acceptors, rotatable bonds, aromatic rings, formal charge, molar refractivity, Lipinski compliance, and molecular complexity. 2. Pharmacokinetic and Drug-Likeness Attributes:Predicted absorption, bioavailability, solubility, blood–brain barrier permeability, plasma protein binding probability, and overall drug-likeness index. 3. Protein Structural and Functional Characteristics:Target-specific attributes such as sequence length, molecular weight, protein family, enzyme class, subcellular localization, domain count, active-site residue count, and binding pocket volume. 4. Ulcerative Colitis Disease-Context Features:Disease relevance indicators derived from gene expression, tissue specificity, inflammatory pathway involvement, cytokine signaling participation, oxidative stress response, immune regulation roles, literature association scores, and biomarker relevance. 5. Molecular Docking Interaction Features:Docking-derived descriptors including binding energy, hydrogen bonding, hydrophobic contacts, electrostatic interactions, π–π stacking, pocket occupancy ratio, and pose stability score. The primary prediction target is a binary label interaction_label, indicating whether berberine is likely to interact with the given protein target. The dataset exhibits mild class imbalance to reflect realistic biological screening conditions. This dataset is not time-series data; instead, it represents independent compound–target interaction instances suitable for classification, regression, graph learning, and explainable AI applications in drug discovery.



