MMTT-Bench
收藏资源简介:
MMTT-Bench是新加坡国立大学构建的合成多模态时间序列与文本注释基准数据集,旨在为信息论指标评估提供真值可控的测试平台。该数据集以正弦波信号为基础,随机插入恒定值平直段(约占30%时间轴),并为每个时间点生成正确、错误和无关三类文本注释。通过完全可控的生成过程,数据集确保了注释与未来信号的真实信息含量已知。它主要用于多模态时间序列预测领域,解决现有实际数据无法验证互信息估计器准确性的关键瓶颈,助力研究者筛选有益文本、评估融合策略,从而提升预测模型的鲁棒性与可解释性。
MMTT-Bench is a synthetic multimodal time series and text annotation benchmark dataset developed by the National University of Singapore, designed to provide a ground-truth-controllable testbed for evaluating information-theoretic metrics. Built upon sinusoidal signals, the dataset incorporates randomly inserted constant-value flat segments that occupy approximately 30% of the total timeline, while generating three categories of text annotations including correct, incorrect and irrelevant for each time point. Through a fully controllable generation process, the dataset ensures that the true information content of both the annotations and future signals is fully known. Primarily used in the field of multimodal time series forecasting, MTTT-Bench addresses the critical bottleneck where existing real-world data cannot verify the accuracy of mutual information estimators, helping researchers screen beneficial text, evaluate fusion strategies and thereby improve the robustness and interpretability of forecasting models.
MMTT-Bench 数据集概述
基本信息
- 数据集名称:MMTT-Bench
- 许可证:CC-BY-4.0
- 语言:英语(en)
- 数据集规模:1K<n<10K
标签
- 多模态(multimodal)
- 融合(fusion)
- 文本(text)
- 时间序列(time-series)
数据配置与划分
- 配置名称:default
- 数据文件划分:
- 训练集(train):
train.json - 验证集(validation):
val.json - 测试集(test):
test.json
- 训练集(train):




