遇见数据集

Dataset for From Literature to Lab LLM powered catalyst synthesis Protocol Generation

收藏
Zenodo2025-09-19 更新2026-05-26 收录
官方服务:

资源简介:

Code are hosted at Github: https://github.com/nsndimt/ChemSynthesis Due to potential copyright issue, please email zhangyue@udel.edu or hfang@udel.edu to access the human annotated dataset. Human Annotated Dataset Files: stage1/stage2 + train/eval/input.json Dataset split suffix meaning: input: all 250 paragraphs train: 200 paragrahs used for training eval: 50 paragraphs used for validation Manual Evaluation Logs Files: stage1_manual_evaluation.csv and stage2_manual_evaluation.csv ID meaning: paragraphs with their ID from 1 to 50 are in the same order as the eval dataset split LLM Checkpoint and Prediction Prediction Files: stage1_pred.tar.gz and stage2_pred.tar.gz Checkpoint File: llamafactory_checkpoint.tar Copy of all LLM checkpoints and predictions used in the paper, give a reference for reproduction Note: all stage 2 prediction are generated by end-to-end testing, meaning we first use Stage 1 LLM to generate predictions and then sent these to Stage 2 LLM Code File: llamafactory_dataset.tar.gz Compressed LLaMA-Factory Dataset files to help people runing Github Codes Github Zip: ChemSynthesis-main.zip Checkpoint containing all github codes and readme files, just keep a second copy of it

提供机构:
Zenodo
创建时间:
2025-09-19
二维码
社区交流群
二维码
科研交流群
商业服务