遇见数据集

Processed Multimodal NMR and MS Spectroscopic Dataset for Set-Structured Representation Learning

收藏
Zenodo2026-06-04 更新2026-06-05 收录
官方服务:

资源简介:

This dataset contains the processed multimodal spectroscopic data (¹H-NMR, ¹³C-NMR, and MS) used in the study on task-dependent value of pretraining in multimodal spectral learning. The raw data originates from the multimodal spectroscopic dataset introduced in Alberts et al. (2024). It has been processed into shard-based tensors where each sample is represented as a set of peak vectors. Each peak is described by a 24-dimensional feature vector, and each sample is padded or truncated to a maximum of 256 peaks. The encoder hidden dimension is 256. Key characteristics: - Splits: Train/Validation/Test with 80/10/10 ratio, random seed 42. - Used for: Masked spectroscopic modeling pretraining of Set Transformer and strict-protocol H/C cross-modal alignment experiments (with leakage-free protocol where H-branch sees only ¹H peaks and C-branch sees only ¹³C peaks). - Accompanying code: Available in the linked GitHub repository (https://github.com/RichardHuang0001/ms-nmr-representation-release) and the Zenodo Software record (DOI 10.5281/zenodo.20519353). This dataset enables full reproducibility of the pretraining and downstream experiments described in the associated paper submitted to Digital Discovery.

提供机构:
Zenodo
创建时间:
2026-06-04
二维码
社区交流群
二维码
科研交流群
商业服务