Dataset for From Literature to Lab LLM powered catalyst synthesis Protocol Generation
收藏资源简介:
Code are hosted at Github: https://github.com/nsndimt/ChemSynthesis Due to potential copyright issue, please email zhangyue@udel.edu or hfang@udel.edu to access the human annotated dataset. Human Annotated Dataset Files: stage1/stage2 + train/eval/input.json Dataset split suffix meaning: input: all 250 paragraphs train: 200 paragrahs used for training eval: 50 paragraphs used for validation Manual Evaluation Logs Files: stage1_manual_evaluation.csv and stage2_manual_evaluation.csv ID meaning: paragraphs with their ID from 1 to 50 are in the same order as the eval dataset split LLM Checkpoint and Prediction Prediction Files: stage1_pred.tar.gz and stage2_pred.tar.gz Checkpoint File: llamafactory_checkpoint.tar Copy of all LLM checkpoints and predictions used in the paper, give a reference for reproduction Note: all stage 2 prediction are generated by end-to-end testing, meaning we first use Stage 1 LLM to generate predictions and then sent these to Stage 2 LLM Code File: llamafactory_dataset.tar.gz Compressed LLaMA-Factory Dataset files to help people runing Github Codes Github Zip: ChemSynthesis-main.zip Checkpoint containing all github codes and readme files, just keep a second copy of it Revision Updates according to reviewer's comments are stored in File: dataset_revision.zip



