JobHop v2
收藏资源简介:
JobHop v2是由根特大学AIDA-IDLab团队构建的大规模公开职业轨迹基准数据集,作为原始JobHop数据集的增强版本。该数据集包含从36.1万份匿名简历中提取的167万条工作经历记录,每条记录均映射至ESCO职业分类体系,并具备季度级时间标注和五级教育水平注释。其构建过程采用基于推理控制的大型语言模型从非结构化多语言简历中进行高质量信息抽取,显著提升了数据质量和模式丰富度。该数据集旨在支持职业路径推荐系统的研究与评估,为解决劳动力市场分析、个性化职业建议及技能缺口预测等实际问题提供高质量基准数据。
JobHop v2 is a large-scale open career trajectory benchmark dataset constructed by the AIDA-IDLab team at Ghent University, serving as an enhanced version of the original JobHop dataset. This dataset encompasses 1.67 million work experience records extracted from 361,000 anonymous resumes, with each record mapped to the ESCO occupational classification system, and equipped with quarterly-level temporal annotations and five-tiered education level annotations. Its construction pipeline leverages inference-controlled large language models to perform high-quality information extraction from unstructured multilingual resumes, substantially improving data quality and pattern richness. This dataset aims to support research and evaluation of career path recommendation systems, providing high-quality benchmark data for addressing practical problems including labor market analysis, personalized career advice, and skill gap prediction.
数据集概述
基本信息
- 数据集名称: JobHop v2
- 发布平台: Hugging Face Hub
- 数据集地址: https://huggingface.co/datasets/aida-ugent/JobHop
- 相关项目: STEP(职业路径推荐模型)和 JobHop v2(信息提取流水线)
数据来源
- 原始简历语料库为私有数据(由弗拉芒公共就业服务局 VDAB 提供的假名化人力资源数据),未包含在公开数据集中。
- 公开数据集中的示例/固定数据均为合成数据。
数据内容
- 基于 LLM 的信息提取流水线,将原始的多语言简历转换为 ESCO 编码的职业轨迹数据。
- 包含数据集构建和评估的代码。
相关论文
- STEP: https://arxiv.org/abs/2607.11722
- JobHop: https://arxiv.org/abs/2607.11715
许可证
- MIT 许可证




