遇见数据集

RAITG: Requirement-Aware Intelligent Test Generator Dataset — 312 Natural-Language Requirements, Generated Test Artefacts, and Reference Applications

收藏
Zenodo2026-06-16 更新2026-05-26 收录
官方服务:

资源简介:

Dataset and reproducibility artefacts for the paper "LLM-Based Test Case Generation from Natural-Language Requirements: A Verified Multi-Domain Empirical Study with Symbolic Mutation Indicators". This release contains: 312 natural-language requirements across commercial web (112), financial services (98), and healthcare (102) domains; 936 per-requirement run logs produced on Anthropic Claude Sonnet 4.6 under three experimental conditions (unverified LLM, framework with verification disabled (ablation), and full RAITG framework); aggregate and per-domain result CSVs; three reference applications used as mutation targets (hr-app 142 lines Python, banking-api 126, fhir-lite 147); the five-element prompt-engineering taxonomy (role, context, task, heuristic, output contract); the four-class rule-based verification calculus (structural, logical, coverage, redundancy); the end-to-end experimental pipeline; and the LaTeX source of the companion paper. Honest methodology caveats: (1) the symbolic mutation indicator reported in the companion paper is regex-based and NOT executable PIT/mutmut; it is comparable between RAITG conditions only and is NOT directly comparable to executable mutation scores reported in the manual-suite literature; (2) the manual-baseline 184.1-hour effort estimate is a literature-cited rate (35 min/requirement) times 312 requirements, NOT a measured wall-clock comparison against matched human authors; (3) the unverified-LLM 0% requirement coverage reflects missing output schema (free-text instead of structured JSON), NOT backend incapability. Headline results: requirement coverage 0% (unverified) -> 100% (RAITG); verification pass rate 73.9% (ablation) -> 95.6% (full RAITG), +21.79 pp absolute with Cohen's d = 0.63 and Wilcoxon p < 0.001 paired n = 312; symbolic mutation indicator 46.7% -> 59.5%. Licence: Creative Commons Attribution 4.0 International (CC BY 4.0). Related datasets: - Software Defect Prediction Dataset (DOI 10.5281/zenodo.19682733) - supports Papers A-C of the same series - Self-Healing Test Automation Dataset (DOI 10.5281/zenodo.19684439, latest v2.0.0 at 10.5281/zenodo.20693970) - supports Paper D Source code: https://github.com/javvadivijayprasad/TestCaseGen_Reserach

提供机构:
Zenodo
创建时间:
2026-05-19
二维码
社区交流群
二维码
科研交流群
商业服务