遇见数据集

Datasets for MulitModalSpectralTransformer project

收藏
Zenodo2025-07-25 更新2026-05-26 收录
官方服务:

资源简介:

Data Repository for "Advancing Structure Elucidation with a Flexible Multi-Spectral AI Model" IMPORTANT NOTE: This is part of a multi-part data repository. Due to the large size of the files, the complete dataset, models, and experimental results are split across three separate Zenodo uploads. To fully reproduce the findings of our publication, you must download all files from all three of the following links: Data Repository Part 1: https://doi.org/10.5281/zenodo.16076914 Data Repository Part 2: https://doi.org/10.5281/zenodo.16257786 Data Repository Part 3: https://doi.org/10.5281/zenodo.16283829 This repository contains the complete dataset, pre-trained models, and experimental results required to reproduce the findings presented in our publication, "Advancing Structure Elucidation with a Flexible Multi-Spectral AI Model." The source code is available separately on GitHub. Contents The downloaded files are split archives that, once combined and extracted, will create the following folders: models/: Pre-trained model weights for MultiModalSpectralTransformer (MMST), SGNN, Mol2Mol, and ChemProp-IR networks. data/: Complete training and validation datasets, including: ZINC dataset (4M molecules with simulated spectra) PubChem dataset (1.5M molecules with simulated spectra) IBM Alberts dataset (650k molecules with simulated spectra) past_experiments/: Complete experimental validation results, benchmarks, and reproducibility data. Setup Instructions Clone the GitHub repository:git clone https://github.com/mpriessner/MultiModalSpectralTransformer.git Download all compressed files: Make sure to download every .partXX file from all three Zenodo links provided above. Place all of them together in the same directory. Combine and Extract: Once all parts are downloaded, use a file archiver that supports split .tar.xz archives (like tar on Linux/macOS or 7-Zip on Windows) to extract the contents. You only need to run the extraction command on the first part of each archive (e.g., data.tar.xz.partaa); the tool will automatically find and combine the other parts. Organize Folders: Move the extracted folders (models, data, and past_experiments) directly into the main repository directory you cloned in step 1. After extraction, your repository structure should contain the necessary models/, data/, and past_experiments/ folders with all the files required for running the notebooks and reproducing the results. For implementation details, usage instructions, and the complete source code, please refer to our Github Reporitory: https://github.com/mpriessner/MultiModalSpectralTransformer

《基于柔性多光谱AI模型推进结构解析》数据集仓库 重要提示:本数据集属于多部分组成的数据集仓库。鉴于文件体积较大,完整数据集、模型与实验结果被拆分至三个独立的Zenodo上传包中。若需完整复现本论文的研究成果,您必须下载以下三个链接中的全部文件: 数据仓库第一部分:https://doi.org/10.5281/zenodo.16076914 数据仓库第二部分:https://doi.org/10.5281/zenodo.16257786 数据仓库第三部分:https://doi.org/10.5281/zenodo.16283829 本仓库包含复现论文《基于柔性多光谱AI模型推进结构解析》中研究成果所需的完整数据集、预训练模型及实验结果。源代码已在GitHub平台单独发布。 内容说明 下载的文件为分卷归档文件,合并并解压后将生成以下文件夹: models/:多模态光谱Transformer(MultiModalSpectralTransformer,MMST)、SGNN、Mol2Mol及ChemProp-IR网络的预训练模型权重。 data/:完整的训练与验证数据集,包含: - ZINC数据集(含模拟光谱的400万分子) - PubChem数据集(含模拟光谱的150万分子) - IBM Alberts数据集(含模拟光谱的65万分子) past_experiments/:存储完整的实验验证结果、基准测试数据与可复现性实验数据。 设置指南 1. 克隆GitHub仓库:执行命令 git clone https://github.com/mpriessner/MultiModalSpectralTransformer.git 2. 下载所有压缩文件:请从上述提供的三个Zenodo链接中下载所有.partXX格式的分卷文件,并将所有文件放置于同一目录下。 3. 合并并解压:下载完成所有分卷后,使用支持分卷.tar.xz归档的解压工具(如Linux/macOS系统下的tar工具,或Windows系统下的7-Zip),仅需对归档的第一部分执行解压命令(例如data.tar.xz.partaa),工具将自动识别并合并其余分卷文件。 4. 整理文件夹:将解压得到的models、data及past_experiments三个文件夹直接移动至步骤1中克隆的主仓库目录下。 解压完成后,您的仓库结构将包含运行Jupyter Notebook与复现研究成果所需的models/、data/及past_experiments/文件夹及全部相关文件。 如需了解实现细节、使用说明与完整源代码,请参阅我们的GitHub仓库:https://github.com/mpriessner/MultiModalSpectralTransformer

提供机构:
Zenodo
创建时间:
2025-07-25
二维码
社区交流群
二维码
科研交流群
商业服务