遇见数据集

RAITG: Requirement-Aware Intelligent Test Generator — Dataset, Pipeline, and Reproducibility Bundle

收藏
Zenodo2026-08-07 更新2026-08-13 收录
官方服务:

资源简介:

Reproducibility bundle for **RAITG**, a 4-stage LLM-based pipeline that generates mutation-adequate test cases from natural-language requirements. ## v3.0.0 (2026-08-07) — What's new - **Executable mutation scoring on the full 362-requirement corpus** (was pilot 28-req in v2.0.0) - **New logistics-app hand-authored Manual baseline** (Threat T10 closed): 22 pytest-style tests, 98.73% kill rate on 79 mutants - **Cross-provider sensitivity study**: Anthropic Claude Sonnet 4.6 (baseline) vs OpenAI GPT-4o-mini on same 362-req corpus and same 273 mutants - Sonnet ablation: 85.71% aggregate kill rate - GPT-4o-mini ablation: 51.65% (−34.06 pp vs Sonnet) - Grounding: Sonnet 98% runnable, GPT-4o-mini 48% - **2-seed variance study** on OpenAI GPT-4o-mini ablation: 1.83 pp aggregate variance (well below RAITG-vs-Manual gaps of 9–14 pp) - **6th prompt element added**: Mutation-Resilience Directive - **Related work expanded**: Meta ACH (Foster et al. 2025), LLMorpheus (Tip et al. 2024), PRIMG (Bouafif et al. 2025) ## Key findings (executable mutation harness, 273 frozen AST-level mutants) **Finding 1** — SUT-source injection lifts runnable fraction from 0% (0 / 298) to 98% (~2 710 / 2 760). +98 pp on a single-variable change. **Finding 2** (Sonnet-conditional) — RAITG's best condition beats hand-authored Manual on 3 of 4 SUTs at ~27× effort reduction: | SUT | Manual | RAITG best | Gap | |---|---|---|---| | banking-api | 86.84% | 96.05% | **+9.21 pp** | | fhir-lite | 75.44% | 89.47% | **+14.03 pp** | | hr-app | 85.25% | 98.36% | **+13.11 pp** | | logistics-app | 98.73% | 67.09% | **−31.64 pp** | Mutant-weighted aggregate: Manual 87.55% vs RAITG-ablation 85.71% (−1.84 pp). **Finding 3** — Rule-based verification is aggregate-neutral (−0.36 pp, within noise) but domain-conditional per-SUT: helps fhir-lite (+5.26 pp), neutral on banking-api, mildly hurts hr-app (−1.64 pp) and logistics-app (−3.80 pp). ## Package contents - `paper/` — Springer ASE PDF (35 pages), LaTeX source, cover letter, 3 figures - `data/` — 362 requirements + 4 SUT sources (`app.py`) - `scripts/` — RAITG pipeline, 4-provider LLM adapters, mutation harness, hand-authored baselines - `results/anthropic_sonnet/` — 1086 Sonnet run logs (362 reqs × 3 conditions) - `results/openai_gpt4o_mini/` — 1086 GPT-4o-mini run logs - `results/seed_variance/` — 362 seed-1 + 362 seed-2 ablation logs - `results/mutation_scores/` — aggregate and per-SUT CSVs - `docs/` — REPRODUCIBILITY.md, FINALIZE_BUNDLE.md **License:** Apache 2.0 **Author:** Vijay Prasad Javvadi, Independent Researcher, Plainsboro NJ USA — ORCID [0009-0004-1192-6906](https://orcid.org/0009-0004-1192-6906)

提供机构:
Zenodo
创建时间:
2026-08-07
二维码
社区交流群
二维码
科研交流群
商业服务