遇见数据集

Metadata and Lucene Index for Evaluating Social Science Integrated Search with MIRA

收藏
Zenodo2026-04-22 更新2026-05-26 收录
官方服务:

资源简介:

Metadata and Lucene index for the MIRA dataset, a novel test collection designed to address the critical evaluation gap in multi-categorical information retrieval. The modern search experience is integrated, yet IR benchmarks have lagged behind, constrained by a lack of collections that mirror this reality. MIRA dataset directly confronts this challenge by providing a unified framework encompassing four distinct scholarly categories - Publications, Research Data, Variables and Instruments & Tools - all grounded in real user queries from the GESIS Search platform. The collection version 1 contains metadata on 7,634 research datasets; 206,434 high-quality metadata variables; 604 instruments & tools; and 254,097 publications with a total of 468,769 documents, provided as a set of JSON files. This dataset includes: The metadata in JSON format The Lucene (version 8) index of the metadata License information Find the software code to use the test collection under https://github.com/suchanadatta/MIRA-LLM-Assisted-Benchmark/tree/main Please cite the publication: Türkmen, M. Deniz, Suchana Datta, Dwaipayan Roy, Daniel Hienert, Philipp Mayr, and Derek Greene. 2026 (Forthcoming). "MIRA: An LLM-Assisted Benchmark for Multi-Category Integrated Retrieval." In SIGIR '26: Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval

提供机构:
Zenodo
创建时间:
2026-04-22
二维码
社区交流群
二维码
科研交流群
商业服务