TechWolf/Skill-extraction-House-graded
收藏资源简介:
skill-extraction-house-graded数据集是一个用于技能提取任务的分级相关性标注数据集。它基于TechWolf/skill-extraction-house中的句子,并针对ESCO v1.1.0技能分类法进行标注。数据集遵循BEIR(信息检索基准)约定,包含三个主要配置:queries(查询,包含句子ID和文本)、corpus(语料库,包含ESCO技能URI、英文首选标签、英文描述和版本)和qrels(查询-文档相关性,包含查询ID、语料ID和0-4的分级分数)。validation分割提供完整的分级相关性标注(0-4分),其中0表示完全不相关,4表示明确相关;test分割目前仅提供二进制相关性标注(0或1),完整分级标注将在后续发布。该数据集用于评估信息检索或技能推荐系统,支持对句子与技能之间相关性的细粒度分析。数据来源包括欧洲委员会的ESCO分类(CC BY 4.0许可)和TechWolf的源句子(同样CC BY 4.0许可)。
The skill-extraction-house-graded dataset is a graded-relevance annotation dataset for skill extraction tasks. It is based on sentences from TechWolf/skill-extraction-house and annotated against the ESCO v1.1.0 skill taxonomy. The dataset follows the BEIR (Benchmarking Information Retrieval) convention and includes three main configurations: queries (containing sentence IDs and texts), corpus (containing ESCO skill URIs, English preferred labels, English descriptions, and version), and qrels (query-document relevance scores with query IDs, corpus IDs, and graded scores from 0 to 4). The validation split provides full graded relevance annotations (0-4), where 0 indicates totally unrelated and 4 indicates explicitly demonstrated relevance; the test split currently offers real but binary relevance labels (0 or 1), with fine-grained graded annotations to be released later. This dataset is designed for evaluating information retrieval or skill recommendation systems, enabling fine-grained analysis of relevance between sentences and skills. Data sources include the European Commissions ESCO classification (licensed under CC BY 4.0) and source sentences from TechWolf (also CC BY 4.0).




