RUST-BENCH
收藏资源简介:
RUST-BENCH是一个包含7,966个问题的数据集,这些数据来自2,031个真实世界的表格,涵盖科学和体育两个领域。与现有的大多数基准测试不同,RUST-BENCH在表格长度、异构性、领域特异性和推理复杂性四个维度上对语言模型进行了评估,提供了一个全面且真实的评估框架。数据集通过一个LLM驱动的混合符号-语义生成管道构建,该管道系统地构建了高质量的、多跳推理问题,这些问题基于真实世界的半结构化表格,同时降低了人工标注成本。RUST-BENCH旨在推动表格推理研究的发展,并为LLM在表格推理方面的研究提供一个挑战性的测试平台。
RUST-BENCH is a dataset consisting of 7,966 questions sourced from 2,031 real-world tables spanning two domains: science and sports. Unlike most existing benchmarks, RUST-BENCH evaluates language models across four dimensions: table length, heterogeneity, domain specificity, and reasoning complexity, providing a comprehensive and realistic evaluation framework. The dataset is constructed via an LLM-driven hybrid symbolic-semantic generation pipeline, which systematically generates high-quality, multi-hop reasoning questions grounded in real-world semi-structured tables while reducing manual annotation costs. RUST-BENCH aims to advance table reasoning research and serve as a challenging testbed for studies on LLMs' table reasoning capabilities.




