Fine-tuning large language models to generate single-atom catalyst synthesis procedures
收藏资源简介:
This repository contains five JSON files that serve as the supplementary dataset for our work, “Fine-tuning large language models to generate single-atom catalyst synthesis procedures.” The first three files constitute the training, validation, and test sets used to fine-tune and evaluate the IBM Granite 3.2-8B Instruct model on single-atom catalyst (SAC) synthesis procedures. These files include 6648 (train), 1663 (validation), and 565 (test) synthesis paragraphs, respectively, along with metadata for the source publications from which the paragraphs were extracted. Two additional files: HT_SP_synthesis_validation and SP_synthesis_validation provide the human validation and scoring of the LLM-generated SAC synthesis recipes as described in the manuscript
本仓库包含五个JSON格式文件,作为我们题为《微调大语言模型(Large Language Model, LLM)以生成单原子催化剂(single-atom catalyst, SAC)合成流程》的研究工作的补充数据集。 前三个文件为用于在单原子催化剂合成流程任务上微调并评估IBM Granite 3.2-8B Instruct模型的训练集、验证集与测试集。这三个文件分别包含6648条训练集合成段落、1663条验证集合成段落及565条测试集合成段落,同时附带了这些段落来源的学术出版物元数据。 另有两个附加文件:HT_SP_synthesis_validation与SP_synthesis_validation,分别提供了本论文手稿中提及的、由大语言模型生成的SAC合成配方的人工验证与评分结果。



