遇见数据集

Replication package for "Vibe Testing for REST APIs: A Context-Taxonomic Benchmark of LLM-Generated Test Suites"

收藏
Mendeley Data2026-07-02 收录
官方服务:

资源简介:

This is the complete replication package for the article "Vibe Testing for REST APIs: A Context-Taxonomic Benchmark of LLM-Generated Test Suites". Vibe Testing is an execution-guided workflow in which a large language model (LLM) synthesizes, runs, diagnoses, and repairs executable REST API test suites in a controlled feedback loop. The study quantifies how the type and amount of supplied context shape both the validity and the defect-detection power of the generated tests, using a five-level context taxonomy (L0: OpenAPI specification only; L1: + business documentation; L2: + source code; L3: + database schema; L4: + reference tests) on a purpose-built, contamination-controlled 50-endpoint Bookstore API. Package contents (folders 01-10): - 01_framework: the Vibe Testing pipeline (context assembly, planning, generation, validation, metrics, repair) and all experiment configurations. - 02_target_api: the 50-endpoint Bookstore REST API under test (FastAPI + SQLite). - 03_context_artifacts: the five context-level input artifacts. - 04_context_axis: balanced experiment (5 cloud LLMs x 5 levels x 5 runs = 125 runs) with per-run metrics and all 125 generated suites. - 05_temperature_axis: temperature sweeps (7 models x 6 temperatures at L1). - 06_openweight_context_matrix: auxiliary open-weight model results. - 07_mutation_testing: 13 hand-curated crud.py mutants and per-suite kill results for all 125 cloud suites. - 08_instruction_following: instruction-following-compliance probe results. - 09_statistics: nonparametric analyses (Friedman, Kendall's W, Wilcoxon, Mann-Whitney, Cliff's delta, bootstrap CIs). - 10_figures: the eight published figures and their generator scripts. MANIFEST.sha256 lists every file with a checksum. All framework and analysis code is in English; the target Bookstore API models a Czech-domain bookstore, so its documentation and a few domain strings are preserved verbatim in the original language as the exact inputs supplied to the models, which does not affect the measured results. Licensing: software components are under the MIT License (LICENSE); data, results, and figures are under CC BY 4.0 (DATA-LICENSE.txt). Citation metadata is in CITATION.cff.

本数据集为论文《面向REST API的Vibe测试:基于大语言模型(LLM)生成测试套件的上下文分类基准》的完整复现包。 Vibe测试(Vibe Testing)是一种执行引导式工作流,大语言模型(LLM)可在受控反馈循环中完成可执行REST API测试套件的合成、运行、诊断与修复操作。本研究依托专为实验构建、且经过污染控制的50端点书店REST API,采用五级上下文分类体系(L0:仅开放API规范;L1:追加业务文档;L2:追加源代码;L3:追加数据库架构;L4:追加参考测试用例),量化了所提供上下文的类型与数量对生成测试用例的有效性及缺陷检测能力的影响。 复现包包含01至10号文件夹,具体内容如下: - 01_framework:Vibe测试流水线(涵盖上下文组装、规划、生成、验证、指标计算与修复模块)及所有实验配置。 - 02_target_api:待测试的50端点书店REST API(基于FastAPI与SQLite构建)。 - 03_context_artifacts:五级上下文层级对应的输入数据集。 - 04_context_axis:平衡实验设置(5个云大语言模型 × 5个上下文层级 × 5次重复实验,共125次运行),包含每次运行的指标数据及全部125套生成的测试套件。 - 05_temperature_axis:温度参数扫描实验(7个模型 × L1层级下的6种温度参数配置)。 - 06_openweight_context_matrix:辅助开源大语言模型的实验结果。 - 07_mutation_testing:针对13个手工编写的crud.py变异体,以及全部125套云模型生成测试套件的变异体杀死率结果。 - 08_instruction_following:指令遵循合规性探测结果。 - 09_statistics:非参数统计分析结果(含Friedman检验、Kendall's W检验、Wilcoxon检验、Mann-Whitney检验、Cliff's delta分析、自助法置信区间)。 - 10_figures:已发表的8幅图表及其生成脚本。 MANIFEST.sha256文件列出了所有文件及其校验和。所有框架与分析代码均采用英文编写;目标书店API建模的是捷克语域的线上书店,因此其文档与少量领域字符串保留了原始语言,作为直接输入至模型的原始数据,该设置不影响实验测量结果。 许可声明:软件组件采用MIT许可协议(详见LICENSE文件);数据、实验结果及图表采用CC BY 4.0许可协议(详见DATA-LICENSE.txt文件)。引用元数据存储于CITATION.cff文件中。

创建时间:
2026-06-25
二维码
社区交流群
二维码
科研交流群
商业服务