m-CHA
收藏资源简介:
m-CHA数据集是由印度理工学院孟买分校的研究团队手工收集的,包含了866个催化meta-C(sp2)-H键激活反应,这些反应源自26篇同行评审的论文。数据集中的反应在底物、偶联伙伴、催化剂、配体、氧化剂、碱和溶剂等方面有所不同。该数据集通过将反应物分子的SMILES字符串连接起来,形成一个适合机器学习模型构建的复合表示。该数据集的应用领域在于优化化学反应的收益率,解决化学反应中的催化剂、反应条件、底物选择等问题。
The m-CHA dataset was manually collected by a research team from the Indian Institute of Technology Bombay, containing 866 catalytic meta-C(sp²)-H bond activation reactions sourced from 26 peer-reviewed papers. The reactions in this dataset vary across substrates, coupling partners, catalysts, ligands, oxidants, bases, solvents and other relevant factors. By concatenating the SMILES strings of the reactant molecules, this dataset constructs a composite representation suitable for the development of machine learning models. The applications of this dataset focus on optimizing chemical reaction yields and addressing challenges such as catalyst selection, reaction condition optimization and substrate selection in chemical reactions.

- 1Efficient Machine Learning Approach for Yield Prediction in Chemical Reactions印度理工学院孟买分校化学系,印度理工学院孟买分校机器智能与数据科学中心 · 2025年



