遇见数据集

sveln/tokamark-dataset

收藏
Hugging Face2026-05-11 更新2026-05-31 收录
官方服务:

资源简介:

TokaMark是一个用于评估AI模型在真实实验数据上的结构化基准数据集,数据来自Mega Ampere Spherical Tokamak (MAST)托卡马克装置。该数据集旨在解决聚变能源研究中数据稀缺、分散和标注不一致的问题,提供多模态异构聚变数据的统一访问、格式协调、元数据标准化、时间对齐和评估协议。数据集包含14个任务,涵盖多种物理机制、诊断工具和目标用例,并提供一个基线模型以促进在统一框架内的透明比较和验证。TokaMark的目标是加速数据驱动的等离子体AI建模进展,支持可持续和稳定聚变能源的实现。数据集完全开源,鼓励社区采用和贡献。

TokaMark is a structured benchmark designed to evaluate AI models on real experimental data collected from the Mega Ampere Spherical Tokamak (MAST). It addresses challenges such as scarce, fragmented, and inconsistently annotated fusion datasets by providing unified access to multi-modal heterogeneous fusion data, harmonizing formats, metadata, temporal alignment, and evaluation protocols. The benchmark includes 14 tasks spanning various physical mechanisms, diagnostics, and target use cases, with a baseline model for transparent comparison and validation within a unified framework. TokaMark aims to accelerate progress in data-driven plasma AI modeling and contribute to achieving sustainable and stable fusion energy. The benchmark, documentation, and tooling are fully open sourced to encourage community adoption and contribution.

提供机构:
sveln
二维码
社区交流群
二维码
科研交流群
商业服务