Dataset for "Large language models in materials science and the need for open-source approaches"
收藏资源简介:
Supporting dataset for: “"Large language models in materials science and the need for open-source approaches”, Fengxu Yang and Jack D. Evans, 2025 This repository provides benchmarking and fine-tuning tools for applying LLMs to metal-organic framework (MOF) synthesis analysis. Key components: Extraction: Performance evaluation of 6 LLMs (GLM-4.5, Qwen series, DeepSeek) on extracting synthesis conditions (precursors, solvents, temperature, time, etc.) from text Prediction: LoRA fine-tuning pipeline for predicting synthesis parameters from unstructured text Datasets: Training and test sets in JSONL format with structured synthesis information Results: Comprehensive benchmarks including accuracy, inference speed, and VRAM requirements Reproduces and extends methodologies from MOF ChemUnity and L2M3 projects.



