遇见数据集

Replication Package: Validation of the Data Related to Vulnerably (Mis)Configured? Exploring 10 Years of Developers' Q&As on Stack Overflow

收藏
Zenodo2026-03-08 更新2026-05-26 收录
官方服务:

资源简介:

Welcome to the public repository for the LLM-based validation (2026) of the paper "Vulnerably (Mis)Configured? Exploring 10 Years of Developers' Q&As on Stack Overflow", accepted at the International Working Conference on Variability Modelling of Software-Intensive Systems (VAMOS) 2024. This repository provides additional information to the conducted exploratory study on configuration-related vulnerabilities (cf. original repository: https://doi.org/10.5281/zenodo.10245332) and the validation file of the dataset (validated by ChatGPT 5.2 and Gemini 2.5 Pro), including the following files: PROMPTS.py: full prompting setup for dataset validation. DATASET_CONFIG_VULN_SO_FULL.csv: sheet containing data of 651 StackOverflow posts, including all metadata and additional classifications based on manual analyses, automatic topic modeling, and LLM-based validation. DATASET_CONFIG_VULN_SO_VALIDATION.csv: sheet containing data of 651 StackOverflow posts, including a reduced set of metadata (used for validation), the classifications of the manual analyses, and LLM-based validation. README.txt Requirements Python >= 3.9; recommended Python packages for analysis: pandas, numpy, scikit-learn, krippendorff, matplotlib No further requirements Further information The dataset is based on a search string (SQL query; August 1, 2023) applied on the Google BigQuery Stack Overflow dataset: ("secur*") AND ("vulnerabilit*" OR "weakness*" OR "breach*" OR "exposure*" OR "CVE*" OR "CWE*") AND ("config*") Originally, the dataset included 1,235 post which were limited by the first and second authors to 651 posts (34 deleted posts, 550 posts out of scope) using the following selection criteria: - The post has been created in the last decade (2013-2022).- The post is still available on the Stack Overflow website.- The post is directly connected to a vulnerability-related issue in the context of configuring. Topic modeling algorithm used: Latent Dirichlet Allocation (LDA)- Settings: 200 iterations (coherence value = 0.6 for k = 7 to 11), α = k, β = 0.01 ChatGPT 5.2 and Gemini were used based on the given prompting strategy License for using the data Creative Commons Attribution 4.0 International The Creative Commons Attribution license allows re-distribution and re-use of a licensed work on the condition that the creator is appropriately credited. Further information: https://creativecommons.org/licenses/by/4.0/legalcode

提供机构:
Zenodo
创建时间:
2026-03-08
二维码
社区交流群
二维码
科研交流群
商业服务