遇见数据集

Replication Package: LLM-Generated Threats to Validity for ICSE Papers

收藏
Zenodo2026-05-30 更新2026-06-05 收录
官方服务:

资源简介:

This replication package contains the data and scripts used to generate and evaluate LLM-generated threats to validity for 375 ICSE research papers. Contents: Input Papers (Input_Papers.zip) - The dataset of 375 ICSE papers used as input for threat generation. Threat Generation Prompt (threat_generation_prompt.txt) - The prompt used to generate threats to validity for all 375 papers. Each paper was passed (excluding its existing threats to validity section) to generate_threats.ipynb alongside this prompt. Generated Threats Dataset (threat_dataset.csv) - The LLM-generated threats for all 375 papers. Note: The dataset contains 110 entries rather than 375 because the generation script occasionally combines multiple papers into a single row when Gemini produces a low number of threats for individual papers. Scripts: paper_split.py: Splits papers (separating the TTV section) for processing. generate_threats.ipynb: Notebook for running the threat generation across all papers. generate_rubric.py: LLM-as-judge script that grades generated threats against a 4-criterion rubric (Relevance, Specificity, Clarity, Mitigation). rubric_count.py: Counts how many threats fall into each scoring bucket per criterion (e.g., how many threats had High Specificity, how many had No Impact for Relevance). compare_threats.py: Semantically compares LLM-generated threats against the paper's actual TTV section using Gemini 2.5 Flash. Buckets each threat into matched, only_AI, or only_Paper across the three validity categories (external, internal, construct). matched threats appear in both the prediction and the paper, only_AI threats were predicted but not in the paper, and only_Paper threats were in the paper but missed by the AI. Generated Rubric (evaluation_rubric.txt) - The rubric used by generate_rubric.py to grade each predicted threat across the four criteria (Relevance, Specificity, Clarity, Mitigation). Scoring Sample (scoring.csv) - A sample of manually graded threats used to establish inter-rater agreement for the rubric criteria.

提供机构:
Zenodo
创建时间:
2026-05-30
二维码
社区交流群
二维码
科研交流群
商业服务