BRIGHT
收藏资源简介:
BRIGHT数据集由香港大学、普林斯顿大学等机构创建,是一个专注于需要深入推理的文本检索基准。该数据集包含1,398条来自经济学、心理学等多个领域的真实查询,数据来源于自然发生或精心策划的人类数据。创建过程中,数据集通过配对真实用户问题与从接受或高票回答中链接的网页来构建。BRIGHT数据集主要应用于评估检索系统在处理复杂查询时的性能,特别是在需要深度理解和上下文关系的情况下。
The BRIGHT dataset, developed by institutions including the University of Hong Kong, Princeton University, and other affiliated organizations, is a text retrieval benchmark dedicated to tasks requiring in-depth reasoning. This dataset comprises 1,398 real-world queries across multiple disciplines such as economics and psychology, with its data originating from either naturally occurring human-generated content or carefully curated human data. During its development, the dataset was constructed by pairing real user questions with webpages linked from accepted or highly-voted answers. The BRIGHT dataset is primarily utilized to evaluate the performance of retrieval systems when handling complex queries, especially in scenarios that demand deep comprehension and contextual relationship understanding.
数据集概述
数据集名称
BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval
作者信息
- Hongjin Su<sup>1</sup>
- Howard Yen<sup>2</sup>
- Mengzhou Xia<sup>2</sup>
- Weijia Shi<sup>3</sup>
- Niklas Muennighoff
- Han-yu Wang<sup>1</sup>
- Haisu Liu<sup>1</sup>
- Quan Shi<sup>2</sup>
- Zachary S. Siegel<sup>2</sup>
- Michael Tang<sup>2</sup>
- Ruoxi Sun<sup>4</sup>
- Jinsung Yoon<sup>4</sup>
- Sercan Ö. Arik<sup>4</sup>
- Danqi Chen<sup>2</sup>
- Tao Yu<sup>1</sup>
机构信息
- <sup>1</sup>The University of Hong Kong
- <sup>2</sup>Princeton University
- <sup>3</sup>University of Washington
- <sup>4</sup>Google Cloud AI Research
资源链接
- 论文: arXiv
- 代码: GitHub
- 数据: Hugging Face

- 1BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive Retrieval香港大学, 普林斯顿大学, 华盛顿大学, Google Cloud AI Research · 2024年



