Train-Ticket as a Microservice Testing Benchmark: A Comparability and Reproducibility Audit
收藏资源简介:
This package accompanies a paper submitted to SANER 2027. It is anonymized for double-blind review: reviewers appear as R1 to R7, and the package contains no author names or affiliations. Files File Content Benchmarking_Train_Ticket_test_generation.xlsx Search exports, screening votes, quality scores, exclusion decisions, and the RQ1 to RQ3 extraction figure.zip Python scripts that regenerate Fig. 2, Fig. 3, and Table II, with the figure PDFs used in the paper Workbook sheets Sheet Content Paper section Inclusion and Exclusion Screening criteria III Search String Final search strings and run dates III, Table I Scopus, IEEE xplore, GS (benchmarking), GS(fault analysis) Raw exports of the four search arms (1153 records) III Abstract & Title Filtering Votes of R1 to R7, disagreement flag, third-reviewer decision (1012 rows, 923 unique records) III full-read (duplicate cleanup) Duplicate consolidation before full-text reading III Full Read v2 Quality scores (QA0 to QA6), exclusion decisions, and the RQ1 to RQ3 extraction for the 75 records sought at full text (21 included) III to VII coding, RQv3 Extraction codebook: fields, allowed values, definitions III Sheets marked (archive) Superseded searches and earlier extraction versions; not used for the reported numbers n/a Bibliographic metadata of screened records (authors, affiliations, abstracts) is kept as exported from the databases. figure.zip figure/ ├── results/ fig_alluvial.pdf Fig. 2 ├── rq1/ make_rq1_fields.py, rq1_counts.py, │ fig_matrix.pdf Fig. 3 ├── rq2/ rq2_counts.py, results_floats.py RQ2 counts, Fig. 2 └── rq3/ rq3_counts.py, table_a_ledger.py RQ3 counts, Table II Reproducing the results Requirements: Python 3.10 or later, with pip install openpyxl matplotlib numpy. Unzip figure.zip next to the workbook, then run the commands from inside figure/. All outputs are written to figure/. cd figure X=../Benchmarking_Train_Ticket_test_generation.xlsx python3 rq3/rq3_counts.py $X # RQ3 counts; writes rq3_extraction.xlsx, rq3_fields.json python3 rq2/rq2_counts.py $X # RQ2 counts, including 0/210 comparable pairs python3 rq2/results_floats.py $X --rq3 rq3_extraction.xlsx # Fig. 2 (fig_alluvial.pdf), results_floats_check.txt python3 rq3/table_a_ledger.py $X rq3_fields.json # Table II (table_a_float.tex) python3 rq1/make_rq1_fields.py $X # writes rq1_fields.json python3 rq1/rq1_counts.py # RQ1 counts, Fig. 3 (fig_matrix.pdf) Each script prints the numbers reported in its Results subsection. Notes Study labels [P1] to [P21] follow the Primary Studies list of the paper. results_floats_check.txt maps each label to its Reference ID in the workbook. Five studies are recoded to Functional Correctness under rule (2) of Section III (RECODE_FC in each script). Artifact links were checked in September 2026. Readiness is judged from documentation and version pinning; no artifact was executed. A language-model-assisted first pass was applied to 105 records of the September 2026 Google Scholar update (sheet scholar2026 (archive)). Every decision was made by a human reviewer. License Data: CC BY 4.0. Code: MIT.



