r-groundbench/mollangbench-edit
收藏资源简介:
该数据集是一个用于化学领域任务的多配置数据集,包含生成任务和视觉问答(VQA)任务。生成任务配置(如generation_easy和generation_hard)涉及分子结构数据,特征包括难度级别、原始SMILES字符串、原始图像、指令、正确SMILES字符串、支架SMILES清洁版本和输出图像。VQA任务配置(如vqa_easy_basic、vqa_hard_advanced等)涉及化学属性问答,特征包括分割类型、属性、原始SMILES字符串、原始图像、指令、问题、正确答案、正确SMILES字符串、属性值、选项(A、B、C、D)的图像和SMILES字符串,以及一些布尔标志(如is_none_of_above)。数据集可能用于训练和评估化学分子生成和视觉问答模型,支持从简单到困难的不同难度级别。每个配置都有训练分割,示例数量较少(每个配置5个示例),总数据集大小从约128KB到378KB不等。
This dataset is a multi-configuration dataset for tasks in the chemistry domain, including generation tasks and visual question answering (VQA) tasks. The generation task configurations (e.g., generation_easy and generation_hard) involve molecular structure data, with features such as difficulty level, original SMILES strings, original images, instructions, correct SMILES strings, scaffold SMILES clean versions, and output images. The VQA task configurations (e.g., vqa_easy_basic, vqa_hard_advanced, etc.) involve chemical property question answering, with features including split type, property, original SMILES strings, original images, instructions, questions, correct answers, correct SMILES strings, property values, images and SMILES strings for options (A, B, C, D), and boolean flags (e.g., is_none_of_above). The dataset is likely used for training and evaluating chemical molecule generation and visual question answering models, supporting different difficulty levels from easy to hard. Each configuration has a train split with a small number of examples (5 examples per configuration), and the total dataset size ranges from approximately 128KB to 378KB.



