CMT-Bench
收藏资源简介:
CMT-Bench数据集由亚利桑那州立大学计算与增强智能学院和Adobe研究院(印度)创建,包含超过6500个板球评论样本及其对应的真实表格,以及一个标签不变的扰动套件。数据集用于测试长上下文文本到表格生成的鲁棒性,要求模型在两个不断发展的表格(击球手和投球手)上进行动态表格生成,受复杂策略(比赛规则)的约束。数据集重点关注长上下文状态跟踪、实体解析和跨跨度聚合,通过三个广泛的维度进行控制鲁棒性探测,以揭示模型的脆弱性。
The CMT-Bench dataset was developed by the School of Computing and Augmented Intelligence at Arizona State University and Adobe Research India. It comprises over 6,500 cricket commentary samples paired with their corresponding ground-truth tables, alongside a label-agnostic perturbation suite. This dataset is designed to test the robustness of long-context text-to-table generation, requiring models to perform dynamic table generation on two evolving tables (batsmen and bowlers) under constraints from complex strategies (cricket match rules). The dataset focuses on long-context state tracking, entity resolution, and cross-span aggregation, and conducts controlled robustness probing across three broad dimensions to uncover model vulnerabilities.

- 1CMT-Bench: Cricket Multi-Table Generation Benchmark for Probing Robustness in Large Language Models亚利桑那州立大学计算与增强智能学院、Adobe研究院(印度) · 2025年



