遇见数据集

Botchuino/OBLIQ-Bench

收藏
Hugging Face2026-05-08 更新2026-05-31 收录
官方服务:

资源简介:

OBLIQ-Bench是一个包含五个检索基准测试的套件,旨在暴露现代搜索系统中的一个盲点:斜向查询(oblique queries),即决定相关性的属性是潜在的,在文档中几乎没有表面表达。相关文档在查询配对时容易识别(推理大型语言模型可以验证它们),但使用任何当前检索系统从大型语料库中检索极其困难。数据集基于三个斜向机制组织:描述性查询(如Twitter-Conflict和WildChat Conversation Errors,查询寻求可以从文档内容推断但过于微妙而无法用当前检索表示表示的潜在属性)、类比查询(如Math Meta-Program和Writing-Style,查询寻求与查询内容共享结构原型的文档,尽管表面主题不同)和舌尖查询(如Congress Hearings,查询将模糊的印象式回忆与特定晦涩文档匹配)。每个任务包括语料库(如推文、对话、数学问题、文本片段、国会听证会段落)、查询和相关判断文件,用于评估检索系统在斜向查询场景下的性能。

OBLIQ-Bench is a suite of five retrieval benchmarks designed to expose a blind spot in modern search systems: oblique queries, where the attributes that determine relevance are latent and have little or no surface expression in the document. Relevant documents are easy to recognize when paired with the query (a reasoning LLM can verify them) but extremely hard to retrieve from a large corpus using any current retrieval system. The dataset is organized by three mechanisms of obliqueness: Descriptive Queries (e.g., Twitter-Conflict and WildChat Conversation Errors, where queries seek a latent property inferable from document content but too nuanced for current retrieval representations), Analogue Queries (e.g., Math Meta-Program and Writing-Style, where queries seek documents sharing a structural archetype with query content despite differing surface topics), and Tip-of-Tongue Queries (e.g., Congress Hearings, where queries match a fuzzy, impressionistic recollection to specific obscure documents). Each task includes a corpus (e.g., tweets, conversations, math problems, text snippets, congressional hearing passages), queries, and relevance judgment files for evaluating retrieval system performance in oblique query scenarios.

提供机构:
Botchuino
二维码
社区交流群
二维码
科研交流群
商业服务