ToRR
收藏资源简介:
ToRR是一个涵盖不同领域和类型的表格推理能力的广泛覆盖的表格基准,包括10个数据集,这些数据集涵盖了从知识提取到文本推理和数值推理不同层次的能力要求。数据集由IBM Research等机构收集,旨在评估模型在表格任务上的性能和鲁棒性。ToRR的设计考虑了模型在不同表格格式下的表现,提供了关于模型在处理表格数据时的鲁棒性测量,对于选择和评估语言模型在表格理解任务上的能力具有重要价值。
ToRR is a comprehensive table benchmark covering table reasoning capabilities across diverse domains and types. It comprises 10 datasets that span capability requirements ranging from knowledge extraction to textual reasoning and numerical reasoning at different levels. Collected by institutions including IBM Research, this benchmark aims to evaluate models' performance and robustness on table-related tasks. Designed to account for model performance across different table formats, ToRR provides robustness measurements for models when processing tabular data, holding significant value for selecting and evaluating language models' capabilities in table understanding tasks.

- 1The Mighty ToRR: A Benchmark for Table Reasoning and RobustnessIBM Research, Bar-Ilan University, Stanford University, MIT · 2025年



