LatEval
收藏资源简介:
LatEval是一个评估大型语言模型横向思维能力的数据集,由清华大学深圳国际研究生院创建。该数据集包含325个高质量样本,来源于多种横向思维谜题网站,涵盖了不完整的故事和真相。创建过程中,通过人工和模型注释筛选和标注关键线索,确保数据集的无害性和挑战性。LatEval旨在通过交互式框架评估模型在提出非传统问题和整合信息以推理真相方面的能力,适用于评估AI助手的横向思维能力。
LatEval is a dataset designed to evaluate the lateral thinking abilities of large language models, developed by the Graduate School at Shenzhen, Tsinghua University. This dataset comprises 325 high-quality samples sourced from multiple lateral thinking puzzle websites, covering incomplete stories and their corresponding truths. During its construction, key clues were screened and annotated through both manual and model-assisted annotation processes to ensure the dataset's harmlessness and appropriate level of challenge. LatEval aims to assess a model's capacity to pose unconventional questions and integrate information for reasoning out the truth via an interactive framework, making it well-suited for evaluating the lateral thinking skills of AI assistants.




