MOFMeld: A Structure–Language Fusion Framework for MOF Property Prediction in Carbon Capture
收藏资源简介:
This Zenodo record accompanies the published article: MOFMeld: a structure–language fusion framework for MOF property prediction in carbon capture Huajie You, Shengde Zhang, Liang Du, Chuxuan Zeng, Teng Zhou and Xiaowen Chu npj Artificial Intelligence, volume 2, Article number: 47 (2026) DOI: https://doi.org/10.1038/s44387-026-00106-1 This Zenodo record contains the processed datasets, model checkpoints, retrieval resources, evaluation outputs, and metadata associated with our study on MOF-specialized language and multimodal modeling for carbon-capture-related reasoning and property prediction. The archived resources support two complementary components of the work: (1) MOFLLaMA, a literature-specialized large language model fine-tuned for MOF question answering, benchmark evaluation, and knowledge-grounded inference; and(2) MOFMeld, a structure-language fusion framework that integrates CHGNet-derived structure embeddings with a frozen MOFLLaMA backbone through a bridge module for structure-aware reasoning and MOF property prediction. The deposit includes the following major resource groups: Baseline_CHGNet: baseline CHGNet checkpoints and prediction outputs for six target properties, used for comparison with MOFMeld. mofbridge_ckpt: MOFMeld bridge checkpoints from stage-I pretraining and stage-II fine-tuning. MOFLLaMA: fine-tuned MOFLLaMA checkpoint files. MOFLLaMA_datasets: supervised fine-tuning and evaluation datasets for MOFLLaMA, including training data, held-out QA data, and easy/hard MCQ benchmark files. MOFLLaMA_KG: knowledge-grounded retrieval resources, including citation metadata, validated knowledge triples, and the FAISS retrieval store used for retrieval-assisted inference. MOFMeld_datasets: datasets, metadata, and evaluation outputs used in MOFMeld, including stage-I pretraining data, stage-II fine-tuning data, hMOF prediction outputs, CoRE-MOF external application files, train/test split metadata, and visualization files. The accompanying GitHub repository provides the released codebase, including preprocessing scripts, training scripts, inference scripts, and runnable demo workflows for both MOFLLaMA and MOFMeld. This Zenodo record is intended for archival distribution of the larger reproducibility assets, including processed datasets, checkpoints, retrieval stores, and result files. Original publisher-provided full-text articles are not redistributed in this archive. Instead, processed derivatives and structured resources are provided, including model-ready QA datasets, retrieval metadata, knowledge triples, structure embeddings, evaluation outputs, and trained checkpoints necessary to reproduce the workflows reported in the manuscript.



