TIME
收藏资源简介:
TIME数据集是一个多层次的综合评估基准,旨在评估大型语言模型(LLMs)在现实场景中的时间推理能力。该数据集由北京大学和华为诺亚方舟实验室的研究人员共同构建,包含38,522个问答对,涵盖三个层次,共11个细粒度的子任务。TIME数据集由三个子数据集组成,分别是TIME-WIKI、TIME-NEWS和TIME-DIAL,分别反映了不同现实世界的挑战。此外,还构建了TIME-LITE,一个高质量的人为标注子集,包含938个精心挑选的实例,以便于未来的研究和标准化评估。
The TIME dataset is a multi-level comprehensive evaluation benchmark designed to evaluate the temporal reasoning abilities of large language models (LLMs) in real-world scenarios. It was jointly constructed by researchers from Peking University and Huawei Noah's Ark Lab. The dataset includes 38,522 question-answer pairs, covering 11 fine-grained subtasks across three hierarchical levels. The TIME dataset consists of three sub-datasets: TIME-WIKI, TIME-NEWS, and TIME-DIAL, which respectively reflect distinct real-world challenges. Additionally, a high-quality human-annotated subset named TIME-LITE has been developed, which contains 938 carefully selected instances to facilitate future research and standardized evaluation.




