granularrxn
收藏资源简介:
GranularRxn 是一个诊断性检索基准数据集,旨在通过三个化学基础维度(抽象对齐、组合对齐和几何对齐)来探究反应嵌入模型的表现。该数据集通过自然语言描述(由 Gemini 2.5 Pro 从反应 SMILES 和 SMARTS 表示生成)来评估 19 种模型的零样本性能,涵盖专有 API、开源编码器、LLM 编码器和领域预训练模型。数据集包含三个任务:任务 1(抽象对齐)测试从具体反应到抽象模板的检索退化;任务 2(组合对齐)测试环基序和反应类约束的联合满足;任务 3(几何对齐)测试区域异构体区分。每个任务包含查询、语料库和相关判断文件,数据规模从数百到数千条记录不等。数据集适用于化学信息检索、嵌入模型评估和相关研究。
GranularRxn is a diagnostic retrieval benchmark dataset designed to evaluate the performance of reaction embedding models across three fundamental chemical dimensions: abstract alignment, compositional alignment, and geometric alignment. This dataset assesses the zero-shot performance of 19 models, including proprietary APIs, open-source encoders, LLM encoders, and domain-pre-trained models, using natural language descriptions generated by Gemini 2.5 Pro from reaction SMILES and SMARTS representations. The dataset contains three tasks: Task 1 (Abstract Alignment) tests retrieval degradation from concrete reactions to abstract templates; Task 2 (Compositional Alignment) tests the joint satisfaction of ring motif and reaction class constraints; Task 3 (Geometric Alignment) tests regioisomer discrimination. Each task includes query, corpus, and relevance judgment files, with data sizes ranging from hundreds to thousands of records. This dataset is applicable to chemoinformatic retrieval, embedding model evaluation, and related research.





