遇见数据集

Dataset for "Large language models in materials science and the need for open-source approaches"

收藏
Zenodo2025-11-07 更新2026-05-26 收录
官方服务:

资源简介:

Supporting dataset for: “"Large language models in materials science and the need for open-source approaches”, Fengxu Yang and Jack D. Evans, 2025 This repository provides benchmarking and fine-tuning tools for applying LLMs to metal-organic framework (MOF) synthesis analysis. Key components: Extraction: Performance evaluation of 6 LLMs (GLM-4.5, Qwen series, DeepSeek) on extracting synthesis conditions (precursors, solvents, temperature, time, etc.) from text Prediction: LoRA fine-tuning pipeline for predicting synthesis parameters from unstructured text Datasets: Training and test sets in JSONL format with structured synthesis information Results: Comprehensive benchmarks including accuracy, inference speed, and VRAM requirements Reproduces and extends methodologies from MOF ChemUnity and L2M3 projects.

本数据集为论文《大型语言模型在材料科学中的应用及开源路径的必要性》(Large language models in materials science and the need for open-source approaches)的配套数据集,作者为杨凤旭与杰克·D·埃文斯,2025年。 本仓库提供了将大语言模型(LLM)应用于金属有机骨架(MOF)合成分析的基准测试与微调工具。 关键组成部分: - 信息抽取:针对6款大语言模型(GLM-4.5、通义千问(Qwen)系列、DeepSeek)开展性能评估,测试其从文本中抽取合成条件(前驱体、溶剂、温度、反应时长等)的能力; - 预测建模:搭建LoRA(低秩适配)微调流水线,实现从非结构化文本中预测合成参数; - 数据集:采用JSONL格式存储的结构化合成信息训练集与测试集; - 测试结果:涵盖准确率、推理速度与显存(VRAM)占用需求的全方位基准测试结果。 本工作复现并拓展了MOF ChemUnity与L2M3项目的研究方法。

提供机构:
Zenodo
创建时间:
2025-11-07
二维码
社区交流群
二维码
科研交流群
商业服务