TechWolf/Skill-normalisation-ESCO-graded
收藏资源简介:
该数据集名为skill-normalisation-esco-graded,是一个用于技能术语标准化和检索的带分级相关性标注的数据集。它基于ESCO v1.1.0技能分类法,将表面技能术语(ESCO替代标签)与ESCO技能分类法进行匹配,并提供了0-4分的分级相关性标注。数据集遵循BEIR标准,适用于MTEB风格的检索评估器。包含三个配置:queries(50个查询,包括查询ID和文本)、corpus(13,891个技能条目,包括技能URI、英文首选标签和描述)和qrels(694,550个相关性标注,包括查询ID、语料库ID和0-4分的评分)。评分越高表示相关性越强,每个查询对应所有ESCO技能条目。数据集分为验证集和测试集,测试集目前提供二元相关性标注(0或1),而细粒度分级标注将在RecSys-HR挑战赛后发布。数据来源于欧盟委员会的ESCO v1.1.0分类,采用CC BY 4.0许可证。
The dataset skill-normalisation-esco-graded provides graded-relevance annotations for surface skill terms (ESCO alt-labels) against the ESCO v1.1.0 skill taxonomy. It follows the BEIR convention for drop-in use with MTEB-style retrieval evaluators. It includes three configs: queries (50 rows with query ID and text), corpus (13,891 rows with skill URI, English preferred label, description, and ESCO version), and qrels (694,550 rows with query-id, corpus-id, and a 0-4 relevance score). Higher scores indicate greater relevance, with every query having one row per ESCO skill. The dataset is split into validation and test sets, where the test set currently has binary relevance labels (0 or 1) derived from public ground truths, while fine-grained graded annotations for the test set are withheld for the RecSys-HR challenge and will be released later. The data is sourced from ESCO v1.1.0 under CC BY 4.0 license.




