Replication package for "Same prompt, different LLM: what survives a model replacement in requirement rewriting"
收藏资源简介:
Replication package for a double-anonymous conference submission. It contains the pipeline, the raw experimental outputs, and the analysis that reproduces every number and figure in the paper. The study compares prompts for rewriting requirements, with and without an evaluator's feedback on the original requirement, on four rewrite LLMs. It asks what carries over when one rewrite LLM replaces another: the ranking of the engineered prompts, their advantage over a baseline prompt, and the value of feedback. Ten prompts run with and without feedback on four LLMs over 200 development requirements. One replacement, GPT-4.1 to GPT-5.6, is repeated on 674 held-out requirements to check the development results. The score that indicates quality comes from an LLM evaluator chosen from sixteen candidates against the judgements of five industrial practitioners, and the main development analyses are repeated with a second evaluator. No API credentials are needed to check the reported numbers. python src/analysis/run_all_analysis.py recomputes every statistic and figure after verifying each raw input artefact against a SHA-256 manifest and checking that no cell is short. See README.md for what to run, and DATA.md for which results trees back which claims.



