Resume Dataset with Generation Pipeline
收藏官方服务:
资源简介:
This repository documents a complete workflow for enhancing and structuring resume data. It starts with raw PDF files (specifically a subset of the Annotated NER PDF Resumes dataset) and generates a highly detailed, machine-readable JSON array suitable for advanced data analysis or machine learning applications. The pdf files were obtained from Mehyaar/Annotated_NER_PDF_Resumes. This process was used to create a structured dataset of only 939 resumes from this dataset.
提供机构:
Zenodo创建时间:
2025-11-18



