JAMP
收藏资源简介:
JAMP是由东京大学创建的日语自然语言推理(NLI)数据集,专注于时间推理。该数据集包含多种时间推理模式,通过精细分析,能够评估语言模型在时间推理上的泛化能力。数据集的创建过程涉及从正式语义测试套件中创建多样化的推理模板,并通过日语案例框架字典和精心设计的模板自动生成多样化的NLI示例,同时控制推理模式和黄金标签的分布。JAMP数据集的应用领域主要在于评估语言模型在时间推理任务上的表现,旨在解决语言模型在处理特定语言现象(如习惯性)时的挑战。
JAMP is a Japanese natural language inference (NLI) dataset focused on temporal reasoning, developed by The University of Tokyo. This dataset covers a variety of temporal reasoning patterns, enabling fine-grained evaluation of language models' generalization abilities in temporal reasoning tasks. The dataset construction process involves creating diverse reasoning templates from formal semantic test suites, and automatically generating varied NLI examples using Japanese case frame dictionaries and meticulously designed templates, while controlling the distribution of reasoning patterns and gold labels. The main application of the JAMP dataset is to evaluate the performance of language models on temporal reasoning tasks, aiming to address the challenges that language models encounter when handling specific linguistic phenomena such as habituality.



