LaRA
收藏资源简介:
LaRA数据集是由香港科技大学、阿里巴巴集团通用人工智能实验室和宾夕法尼亚州立大学共同创建的,包含2326个测试案例,跨越四个实际问答任务类别,涵盖三种自然发生的长文本类型。该数据集旨在为评估长文本处理能力提供严格的基准,包含小说、学术论文和财务报表等不同写作风格和信息密度的长文本。LaRA的任务设计考虑到定位信息、比较文本不同部分、内容推理和检测虚构内容等方面,以全面评估LC LLMs和RAG的能力。
The LaRA dataset was jointly created by The Hong Kong University of Science and Technology, Alibaba Group's General Artificial Intelligence Laboratory, and Pennsylvania State University. It consists of 2326 test cases across four practical question answering task categories, covering three types of naturally occurring long texts. Designed to provide a rigorous benchmark for evaluating long-text processing capabilities, the dataset includes long texts with varying writing styles and information densities such as novels, academic papers, and financial statements. The task design of LaRA takes into account aspects including information localization, comparison of different parts of the text, content reasoning, and detection of fictional content, so as to comprehensively evaluate the capabilities of LC LLMs and RAG.




