CROWN: Curated Repository Of Well-resolved Non-covalent interactions
收藏资源简介:
The development of machine learning models for protein–ligand interactions is fun-damentally constrained by the quality and diversity of available structural data. Ex-isting databases of protein–ligand complexes present researchers with an unsatisfyingtrade-off: carefully curated collections such as PDBBind and HiQBind offer high struc-tural reliability but cover only a narrow slice of the Protein Data Bank (PDB), whilelarge-scale resources like PLInder provide broad coverage at the expense of rigorousquality control. Here, we introduce CROWN (Curated Repository Of Well-resolvedNon-covalent interactions), a machine learning–ready dataset that reconciles this ten-sion by applying a comprehensive, fully automated preprocessing pipeline to the PLIn-der database. Starting from 649,915 protein–ligand interaction systems, CROWNapplies a series of interleaved quality filters and processing stages addressing crys-tallographic resolution, ligand identity, pocket completeness, structural repair, inter-action quality, and protonation at physiological pH. A distinguishing feature of thepipeline is a final constrained energy minimisation step using custom flat-bottomedrestraints, which balances crystallographic evidence with relaxation of intramolecularstrain. This step — absent from existing protein–ligand datasets — produces struc-turally uniform complexes by reconciling the heterogeneous refinement practices of dif-ferent crystallographers and structure determination protocols, without distorting theexperimentally observed binding geometry. The resulting dataset of 153,005 complexesrepresents a roughly four-fold increase in protein and species diversity over PDBBindand HiQBind, while maintaining rigorous structural standards. Importantly, CROWNadopts a geometry-centric design philosophy that treats the 3D arrangement of atomsat the binding interface as a self-consistent source of information, rather than relyingon externally measured binding affinities that cover only a fraction of known struc-tures and introduce well-documented biases. We anticipate that CROWN will serveas a broadly useful resource for training generative models of protein–ligand bindingposes, developing scoring functions, and benchmarking interaction prediction methods.



