Educational Resource Search Relevance Dataset
收藏资源简介:
本研究构建的教育资源搜索相关性数据集,由德国汉诺威莱布尼茨大学L3S研究中心和汉诺威莱布尼茨信息技术中心共同完成。数据集收集了12名参与者在执行19个课程规划相关任务时的401个明确的相关性判断,这些判断基于他们与搜索任务相关的文档互动的思考过程。数据集中的文档涵盖了从网络页面到PDF文档,经过预处理后用于大规模语言模型的相关性评估。该数据集旨在解决教育资源搜索中自动评估相关性的问题,并为大规模语言模型在特定领域搜索中的评估提供了实用的框架。
The educational resource search relevance dataset constructed in this study was jointly developed by the L3S Research Center and the Leibniz Institute of Information Technology Hannover at Leibniz University Hannover, Germany. The dataset contains 401 explicit relevance judgments collected from 12 participants while they completed 19 course planning-related tasks, with these judgments derived from their think-aloud protocols during interactions with documents relevant to the search tasks. The documents in the dataset range from web pages to PDF documents, and have been preprocessed for relevance evaluation with large language models. This dataset aims to address the challenge of automatic relevance assessment in educational resource search, and provides a practical framework for evaluating large language models in domain-specific search scenarios.

- 1Validating LLM-Generated Relevance Labels for Educational Resource Search德国汉诺威莱布尼茨大学L3S研究中心, 德国汉诺威莱布尼茨信息技术中心 · 2025年



