CUP Dataset
收藏资源简介:
CUP数据集是由希腊研究与技术基金会和克里特大学出版社联合构建的希腊语图书检索基准,包含868条涵盖科学、人文与艺术等主题的图书元数据记录,每条记录包含标题、作者、描述及目录等结构化与自由文本字段,其中描述字段平均约225词,目录字段平均226词且长度变异较大。该数据集通过专家人工标注构建了104条多样化查询及其分级相关性判断,旨在评估稀疏、稠密及混合检索方法在希腊语这一形态丰富低资源语言中的性能,特别针对词汇、语义、噪声及跨语言查询场景,为希腊语信息检索研究提供了现实世界的评估基础。
CUP Dataset is a Greek-language book retrieval benchmark jointly constructed by the Foundation for Research & Technology – Hellas and University of Crete Press. It contains 868 book metadata records covering topics including science, humanities and arts. Each record includes structured and free-text fields such as title, author, description and table of contents, where the description field averages approximately 225 words, while the table of contents field averages 226 words with considerable length variability. The dataset includes 104 diversified queries paired with graded relevance judgments developed via expert manual annotation, aiming to evaluate the performance of sparse, dense and hybrid retrieval methods in morphologically rich low-resource languages such as Greek. It specifically targets lexical, semantic, noisy and cross-lingual query scenarios, providing a real-world evaluation foundation for Greek-language information retrieval research.




