lukaskim/ChEMBL-36
收藏资源简介:
ChEMBL 36数据集是ChEMBL数据库(EMBL-EBI)的HuggingFace格式转换版本,ChEMBL数据库是一个手动策划的包含具有药物类似特性的生物活性分子的数据库。该数据集包含三个主要配置:molecules(约240万行,提供所有化合物的规范SMILES表示及相关化学和药物特性,如分子量、LogP、氢键供体/受体数量等)、targets(约1.5万行,包含蛋白质靶标的氨基酸序列、生物分类和蛋白质家族信息)和molecule_target_pairs(约100万到500万行,记录化合物与蛋白质靶标之间的生物活性配对数据,包括标准化活性值和实验类型)。数据集适用于化学、药物发现、生物学等领域的研究,支持SMILES、蛋白质序列和生物活性分析。数据来源于ChEMBL 36 SQLite发布版,采用CC BY-SA 4.0许可证。
ChEMBL 36 is a dataset converted to HuggingFace format from the ChEMBL database (EMBL-EBI), a manually curated database of bioactive molecules with drug-like properties. It includes three configurations: molecules (approximately 2.4 million rows, providing canonical SMILES representations and related chemical and drug properties such as molecular weight, LogP, hydrogen bond donor/acceptor counts), targets (approximately 15,000 rows, containing amino-acid sequences, biological classification, and protein family information for protein targets), and molecule_target_pairs (approximately 1 to 5 million rows, recording bioactivity pairs linking compounds to protein targets, including standardized activity values and assay types). The dataset is suitable for research in chemistry, drug discovery, biology, and other fields, supporting analysis of SMILES, protein sequences, and bioactivity. Data is sourced from the ChEMBL 36 SQLite release and licensed under CC BY-SA 4.0.



