遇见数据集

DORIS-MAE-v0

收藏
Zenodo2023-06-13 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

In scientific research, the ability to effectively retrieve relevant documents based on complex, multifaceted queries is critical. Existing evaluation datasets for this task are limited, primarily due to the high costs and effort required to annotate resources that effectively represent complex queries. To address this, we propose a novel task, <strong>S</strong>cientific <strong>DO</strong>cument <strong>R</strong>etrieval using <strong>M</strong>ulti-level <strong>A</strong>spect-based qu<strong>E</strong>ries (DORIS-MAE), which is designed to handle the complex nature of user queries in scientific research. The DORIS-MAE dataset is publicly available at https://github.com/Real-Doris-Mae/Doris-Mae-Dataset. This dataset is comprised of four main sub-datasets, each serving distinct purposes. The <strong>Query</strong> dataset contains 50 human-crafted complex queries spanning across five categories: ML, NLP, CV, AI, and Composite. Each category has 10 associated queries. Queries are broken down into aspects (ranging from 3 to 9 per query) and sub-aspects (from 0 to 6 per aspect, with 0 signifying no further breakdown required). For each query, a corresponding candidate pool of relevant paper abstracts, ranging from 99 to 130, is provided. The <strong>Corpus</strong> dataset is composed of 363,133 abstracts from computer science papers, published between 2011-2021, and sourced from arXiv. Each entry includes title, original abstract, URL, primary and secondary categories, as well as citation information retrieved from Semantic Scholar. A masked version of each abstract is also provided, facilitating the automated creation of queries. The <strong>Annotation</strong> dataset includes generated annotations for all 83,591 question pairs, each comprising an aspect/sub-aspect and a corresponding paper abstract from the query's candidate pool. It includes the original text generated by ChatGPT (version chatgpt-3.5-turbo-0301) explaining its decision-making process, along with a three-level relevance score (e.g., 0,1,2) representing ChatGPT's final decision. Finally, the <strong>Test Se</strong>t dataset contains human annotations for a random selection of 250 question pairs used in hypothesis testing. It includes each of the three human annotators' final decisions, recorded as a three-level relevance score (e.g., 0,1,2).

提供机构:
Wang, Doris Mae
创建时间:
2023-06-13
二维码
社区交流群
二维码
科研交流群
商业服务