TRAM
收藏资源简介:
TRAM是一个综合的时序推理基准,由斯坦福大学创建,包含10个数据集,旨在全面评估大型语言模型(LLMs)的时序推理能力。这些数据集覆盖了从基础的时序理解到高级的时序解释和计算等多个方面,如事件顺序、持续时间、频率和时序算术等。每个数据集都经过精心设计,以评估模型在不同难度和理解层次上的表现。TRAM不仅包括现有的自然语言理解数据集,还融合了人工编制的模板和问题,以及网络资源和程序生成,总计包含526,068个问题。这些问题通过专家注释和程序生成相结合的方式得出,旨在推动LLMs在理解和推理时间方面的进一步发展,解决复杂叙述和事件因果关系中的时序问题。
TRAM is a comprehensive temporal reasoning benchmark developed by Stanford University, which comprises 10 datasets and aims to comprehensively evaluate the temporal reasoning capabilities of Large Language Models (LLMs). These datasets cover multiple dimensions ranging from basic temporal understanding to advanced temporal interpretation and computation, such as event ordering, duration, frequency, and temporal arithmetic. Each dataset is meticulously designed to assess model performance across varying difficulty levels and comprehension tiers. TRAM not only includes existing natural language understanding datasets, but also integrates manually curated templates and questions, as well as web resources and program-generated content, with a total of 526,068 questions. These questions are derived through a combination of expert annotation and program generation, with the goal of further advancing the development of LLMs in temporal understanding and reasoning, and resolving temporal issues within complex narratives and event causal relationships.

- 1TRAM: Benchmarking Temporal Reasoning for Large Language Models斯坦福大学 · 2023年



