遇见数据集

AI-era software-engineering job postings: a 2024–2026 panel with LLM-derived role, skill, and seniority labels

收藏
Zenodo2026-05-21 更新2026-05-26 收录
官方服务:

资源简介:

Zenodo description (paste into the Description field) This is the Zenodo form description text, kept here for easy re-paste if the form is reset or a v2 release needs the same body. Not part of the release contents — lives at release/ rather than release/staging/. Overview A unified, analysis-ready panel of 114,954 US software-engineering and matched-control job postings drawn from three LinkedIn sources: two 2024 Kaggle snapshots (Arsh Koneru's LinkedIn Job Postings (2023–2024) and asaniczka's 1.3M LinkedIn Jobs & Skills (2024)) and a 2026 first-party scrape across 26 US metropolitan areas. The dataset accompanies a paper on AI-driven restructuring of software-engineering roles. Every posting carries LLM-derived labels for seniority, years-of-experience floor, ghost-job assessment, an 8-enum skill-theme axis (people management, orchestration, verification, mentorship, performance, process scaffolding, legacy stack, context infrastructure), and a 17-enum role-family axis (frontend, backend, ML, AI/LLM engineer, devops, security, QA, and others). The frozen production prompts are included verbatim in the release. Each posting with cleaned text also carries a 3072-dimensional text-embedding-3-large embedding computed over the title plus the boilerplate-removed description core. What's included data/unified_core.parquet — the canonical analysis file, 114,954 rows × 35 columns. data/unified_core_observations.parquet — daily panel (one row per posting × scrape-date) for posting-duration work. scraped_raw/ — 363,060-posting near-raw fallback for the 2026 scrape, joinable by uid for researchers who want to redo cohort/preprocessing choices from scratch. prompts/ — the three frozen production LLM prompts (Stage 9 extraction, Stage 10 classification, Stage 12 skill-theme × role-family). scripts/rejoin_kaggle.py — restores raw 2024 descriptions byte-deterministically from upstream Kaggle source files. CODEBOOK.md, DATASHEET.md, ATTRIBUTIONS.md, croissant.json — full documentation, Gebru-style datasheet, license attributions, and MLCommons Croissant 1.0 metadata. Cohort definition Rows are the intersection of (a) the pipeline's deterministic balanced Stage-9 LLM frame and (b) a confirmed cohort label (LLM-confirmed SWE-or-adjacent or rule-based control). The canonical disjoint analysis frame is is_swe AND NOT is_control (59,954 SWE rows) versus is_control AND NOT is_swe (54,835 control rows); 165 overlap rows are kept in the file but typically excluded. 2024 description policy The raw description column is null on rows where source ∈ {kaggle_arshkon, kaggle_asaniczka}. This avoids redistributing substantively copyrightable text from the upstream Kaggle datasets while preserving every derived column (cleaned text, embedding, all LLM labels) for all rows. The bundled rejoin script restores raw descriptions byte-identically from locally-downloaded upstream Kaggle files; see ATTRIBUTIONS.md for the derivative-work position and any downstream license obligations. License and reuse Released under CC BY 4.0 for data and documentation, MIT for the bundled scripts. Upstream attributions: Arsh Koneru (CC BY-SA 4.0) and asaniczka (ODC-By 1.0). The 2026 scrape is the depositor's first-party collection. If you use this dataset, please cite this Zenodo record. Start with README.md, then CODEBOOK.md for column-level documentation.

提供机构:
Zenodo
创建时间:
2026-05-21
二维码
社区交流群
二维码
科研交流群
商业服务