HAMLET
收藏资源简介:
HAMLET是一个全面的自动化评估框架,用于评估大型语言模型在长上下文中的理解能力。该框架将源文本结构化为三层关键事实层次结构,并采用查询式摘要来评估模型在不同层次上对信息的回忆和忠实表示。HAMLET框架包括三个阶段:查询构建、摘要生成和自动摘要评估。该框架使用16部小说作为数据集,每部小说平均长度为101K tokens,涵盖了从全局主题到具体细节的多层次内容。数据集通过GPT-4o生成关键事实树,并通过查询式摘要任务评估LLMs的回忆和忠实性。
HAMLET is a comprehensive automated evaluation framework designed to assess the long-context understanding capabilities of large language models. This framework structures source texts into a three-tiered key fact hierarchy, and adopts query-based summarization to evaluate models' recall and faithful representation of information across different levels. The HAMLET framework comprises three stages: query construction, summarization generation, and automatic summarization evaluation. The dataset utilized by this framework includes 16 novels, each averaging 101K tokens in length, covering multi-level content ranging from global themes to specific details. Key fact trees for the dataset are generated via GPT-4o, and the recall and faithfulness of LLMs are evaluated through query-based summarization tasks.

- 1通过韩国科学技术院 · 2025年



