遇见数据集

Resume Dataset with Generation Pipeline

收藏
Zenodo2025-11-18 更新2026-05-26 收录
官方服务:

资源简介:

This repository documents a complete workflow for enhancing and structuring resume data. It starts with raw PDF files (specifically a subset of the Annotated NER PDF Resumes dataset) and generates a highly detailed, machine-readable JSON array suitable for advanced data analysis or machine learning applications. The pdf files were obtained from Mehyaar/Annotated_NER_PDF_Resumes. This process was used to create a structured dataset of only 939 resumes from this dataset.

提供机构:
Zenodo
创建时间:
2025-11-18
二维码
社区交流群
二维码
科研交流群
商业服务