DuReaderretrieval
收藏资源简介:
DuReaderretrieval是一个大规模的中文数据集,用于从网络搜索引擎中检索段落。该数据集包含超过90K查询和超过8M独特段落,所有查询均来自百度搜索的真实用户请求,文档段落来自搜索结果。数据集通过远监督和人工标注相结合的方式创建,旨在解决段落检索中的挑战,如显著短语不匹配和语法不匹配。此外,数据集还提供了跨领域和跨语言的评估集,以评估模型的泛化能力和跨语言检索能力。
DuReaderretrieval is a large-scale Chinese dataset for paragraph retrieval from web search engines. It contains over 90K queries and over 8M unique paragraphs, where all queries are real user requests from Baidu Search, and the document paragraphs are sourced from search results. The dataset is developed through a hybrid approach combining distant supervision and manual annotation, aiming to tackle core challenges in paragraph retrieval such as prominent phrase mismatches and grammatical mismatches. In addition, the dataset provides cross-domain and cross-language evaluation sets to assess a model's generalization capability and cross-language retrieval performance.

- 1DuReader_retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search Engine百度公司 · 2022年



