Data Files for TIMS-Bench Towards community standards for benchmarking untargeted trapped ion mobility metabolomics tools and datasets
收藏资源简介:
TitleTIMS-Bench datasets and benchmarking workflow outputs for untargeted trapped ion mobility metabolomics Description / AbstractThis deposition contains the processed benchmarking datasets and reference files used in TIMS-Bench, a reproducible framework for evaluating untargeted TIMS metabolomics software across harmonization, annotation, and downstream performance metrics. The release includes harmonized feature tables, annotation outputs (cosine and spectral entropy), DreaMS-compatible embeddings, and curated ground-truth and library resources. The dataset supports cross-tool benchmarking for MetaboScape, MS-DIAL, and MZmine, and is organized to enable direct reuse of the analysis notebooks and workflow components described in the associated manuscript (Rajkumar et al., 2026). Process Summary (How the data were generated and organized) Raw tool exports were harmonized into a common schema (feature metadata, spectral fields, and sample intensity columns). Harmonized files were annotated using multiple similarity methods, including cosine and spectral entropy workflows. Embedding-based resources were generated/collected for DreaMS-based analyses. Ground-truth benchmark datasets were assembled for ReFRAME, NIST SRM, and plant spike-in evaluations. Reference library files were prepared to support annotation and false-positive structural similarity analyses. More details can be found in the GitHub repository: https://github.com/enveda/tims-bench Data Contents Top-level data groups: cache public_dataset groundtruth_dataset library_spectra Public datasets included (10): MSV000084402 MSV000090327 MSV000091642 MSV000095813 MSV000096189 MSV000096291 MSV000097015 MSV000097967 MTBLS12332 ST002402 Ground-truth datasets included (3): MSV000098263 NIST_SRM plant_spikein Reference/library files included: all_sorted_library_spectra.parquet (about 2.2 GB) nist_srm_spikein_lib.pq plant_spikein_lib.pq reframe_ms2s_with_ccs.parquet reframe_spikein_lib.pq reframe_smiles_list.csv Common per-dataset structure: raw (original tool exports) harmonized (standardized parquet outputs) annotated_cosine_similarity annotated_spectral_entropy annotated_dreams_similarity (where available) embeddings Intended ReuseThis deposition is intended for: Reproducible benchmarking of untargeted TIMS metabolomics tools Method comparison across annotation strategies Development and validation of new harmonizers and scoring methods False-positive and structural-similarity analyses against curated ground truth



