Fine-tuning large language models to generate single-atom catalyst synthesis procedures
收藏资源简介:
This repository contains five JSON files that serve as the supplementary dataset for our work, “Fine-tuning large language models to generate single-atom catalyst synthesis procedures.” The first three files constitute the training, validation, and test sets used to fine-tune and evaluate the IBM Granite 3.2-8B Instruct model on single-atom catalyst (SAC) synthesis procedures. These files include 6648 (train), 1663 (validation), and 565 (test) synthesis paragraphs, respectively, along with metadata for the source publications from which the paragraphs were extracted. Two additional files: HT_SP_synthesis_validation and SP_synthesis_validation provide the human validation and scoring of the LLM-generated SAC synthesis recipes as described in the manuscript



