RULEARENA
收藏资源简介:
RULEARENA是一个新颖且具有挑战性的基准数据集,旨在评估大型语言模型(LLMs)在复杂、现实世界规则下的推理能力。该数据集涵盖了航空行李费、NBA交易和税务政策三个实际领域,包含95条常用且中等复杂的规则和816个测试问题。数据集的创建过程包括从真实场景中收集规则,并构建具有不同难度级别的测试问题。RULEARENA的应用领域主要集中在评估LLMs在实际应用中的规则遵循和推理能力,旨在解决LLMs在复杂规则下的推理和计算问题。
RULEARENA is a novel and challenging benchmark dataset designed to evaluate the reasoning capabilities of Large Language Models (LLMs) under complex, real-world rules. This dataset covers three practical domains: airline baggage fees, NBA transactions, and tax policies, containing 95 commonly used moderately complex rules and 816 test questions. The dataset creation process includes collecting rules from real-world scenarios and constructing test questions with varying difficulty levels. The primary application scenarios of RULEARENA focus on evaluating the rule-following and reasoning abilities of LLMs in practical applications, aiming to solve the reasoning and computational problems faced by LLMs under complex rules.




