Online Appendix for Show Your Title! A Scoping Review on Verbalization in Software Engineering with LLM-Assisted Screening
收藏资源简介:
Validation & Replication Appendix This appendix accompanies "Show Your Title! A Scoping Review on Verbalization in Software Engineering with LLM-Assisted Screening" and contains all artifacts and data necessary to replicate and validate our scoping review methodology. Keywords: verbalization-based techniques · human-centered SE · cognitive aspects of SE · meta-research · qualitative research in SE · knowledge representation in SE · software development practices Overview & Validation Focus This appendix serves a dual validation purpose: enabling complete methodological replication while demonstrating the reliability and consistency of our LLM-assisted approach. All files support our claims about GPT's 91%-92% internal consistency and acceptable 13% disagreement rate between GPT and human reviewers. File Descriptions Keyword Files (cf. se_keywords.json and psy_keywords.json)- `se_keywords.json` - Software Engineering keywords organized by inclusion question tags - `psy_keywords.json` - Psychology keywords organized by inclusion question tags Purpose: Complete list of keywords used for literature search. Essential for replicating our search strategy that operationalizes the seven inclusion questions. Venue Information (cf. key venues.md)- `key venues.md` - Selected venues per field used in our literature search Purpose: Documents the publication venues that define disciplinary boundaries for our intersection analysis. GPT Processing Materials (cf. multi paper relevancy prompt for OpenAI.md and asses_relevancy_with_gpt_by_title.py)- `multi paper relevancy prompt for OpenAI.md` - Standardized prompt used for GPT-4.1 relevance assessment- `asses_relevancy_with_gpt_by_title.py` - Python script for conducting batch relevance assessments using GPT Purpose: Enable exact replication of our LLM-assisted screening methodology using title-based assessment in batches of 10 papers. Results and Analysis Files (cf. relevant_justification_grouping_by_GPT_runX.json and tags_distribution.json)- `relevant_justification_grouping_by_GPT_runX.json` (where X = 1, 2, or 3) - GPT-generated thematic grouping results across three independent runs, each containing: - `label`: Theme name - `reasoning_pattern`: Informal theme definition - `prominence_reason`: Explanation of theme prominence - `justifications`: Array of example justifications - `tags_distribution.json` - Statistical distribution of inclusion question tags across relevant papers, supporting Figure 2 visualization Purpose: Support our thematic analysis and enable replication of the prominent theme identification process. Validation Data (cf. manual_validation_assessment_agreement.json and psy_on_se\GPT_relevancy_assessment_agreement.json)- `manual_validation_assessment_agreement.json` - Human reviewer agreement data showing perfect agreement achieved in 60% of cases, with 40% showing partial disagreements - `psy_on_se\GPT_relevancy_assessment_agreement.json` and `se_on_py\GPT_relevancy_assessment_agreement.json` - GPT internal consistency data demonstrating 91%-92% self-agreement across multiple assessments Purpose: Validate our claims about GPT consistency exceeding human reviewer agreement and support the 13% final disagreement rate between GPT and human judgments. Scripts Note that some scripts might need some minor editing, so the included paths match your file system.Some of the data were manually aggregated by running the same scripts multiply times.- src - all the scripts we used to produce the above mentioned raw data- pyproject.toml - project metadata, this files makes it easier to install the dependencies Key Statistics Supported by This Data - 9,265 papers processed after deduplication (5,386 SE on PSY + 3,879 PSY on SE)- 1,675 papers marked as relevant through majority voting across three GPT runs- 792 papers (14.7%) marked relevant in SE on PSY intersection - 883 papers (22.76%) marked relevant in PSY on SE intersection- 100-paper validation sample with 95% confidence level and margin of error below 10%- Sample composition: 42% PSY on SE papers, 58% SE on PSY papers- Relevance distribution in sample: PSY on SE (23% relevant, 77% irrelevant), SE on PSY (20% relevant, 80% irrelevant) Validation Results - GPT internal consistency: 91%-92% perfect self-agreement - Human reviewer agreement: 60% perfect agreement, 40% partial disagreements- GPT vs Human disagreement: 13% final disagreement rate after consolidation- GPT divergence rates: 9%-8% partial divergence for SE on PSY and PSY on SE respectively Technical Requirements for Replication - Scopus API access for literature search replication- GPT-4.1 (2025-05-15) for exact prompt execution- Python environment for script execution- JSON processing capabilities for data analysis This appendix enables full replication of our methodology while validating the effectiveness of LLM-assisted screening in interdisciplinary scoping reviews.



