TOMG-Bench
收藏资源简介:
TOMG-Bench是由香港理工大学、上海交通大学和上海人工智能实验室联合创建的文本化开放分子生成基准数据集,旨在评估大型语言模型在分子生成任务中的能力。该数据集包含三个主要任务:分子编辑、分子优化和定制分子生成,每个任务包含三个子任务,每个子任务有5000个测试样本,总计45000条数据。数据集通过随机采样和化学工具箱RDKit进行构建,确保分子生成的准确性和有效性。TOMG-Bench的应用领域主要集中在药物发现和材料科学,旨在解决传统分子生成方法的局限性,推动分子设计领域的创新。
TOMG-Bench is a text-based open molecular generation benchmark dataset jointly created by The Hong Kong Polytechnic University, Shanghai Jiao Tong University, and Shanghai AI Laboratory, which aims to evaluate the capabilities of large language models (LLMs) in molecular generation tasks. This dataset encompasses three core tasks: molecular editing, molecular optimization, and customized molecular generation. Each task consists of three subtasks, with 5000 test samples per subtask, totaling 45,000 data entries. The dataset is constructed via random sampling and the chemical toolbox RDKit, ensuring the accuracy and validity of the generated molecular structures. The application domains of TOMG-Bench primarily focus on drug discovery and materials science, with the objectives of addressing the limitations of traditional molecular generation methods and promoting innovation in the field of molecular design.




