遇见数据集

Artifact for "Remembered, Not Generated: A Renamed-Schema Control for LLM-Provisioned Test Data, the Cost of Planning Instead of Emitting, and What Valid Data Does Not Do for LLM-Generated Tests"

收藏
Zenodo2026-09-24 更新2026-10-01 收录
官方服务:

资源简介:

Complete artifact for the empirical study of LLM test-data provisioning submitted to the Journal of Systems and Software: (E1) direct emission vs plan-then-execute (SynthData 1.2.2) vs no-LLM vs Faker vs human fixtures on five open-source PostgreSQL schemas (Pagila, Chinook, Northwind, employees, Dell DVD Store 2), two providers (Claude Sonnet 4.6, GPT-4o), two runs, scored by row-by-row insertion into PostgreSQL 16 with all constraints on and by column-wise fidelity metrics against the human data; (E1b) the same LLM arms on identifier-renamed schemas as a memorisation control; (E2) the provisioned fixtures injected into the RAITG LLM test-generation pipeline (172 requirements, 4 services, 293 frozen mutants) in three arms plus (E2c) a pre-registered state-contract arm, with executable mutation scoring. Contains the pre-registered research plan and the E2c pre-statement, frozen inputs with checksums (canonical and renamed DDLs, business cases, row targets, rename maps, E2 DDLs/cases/fixtures/prompt blocks), every LLM cell's output with per-call token and latency logs, all score and metric files, per-mutant kill vectors, known-good lists and pre-screen records, the scoring harness, the analysis scripts (stats_paper.py prints every statistic in the paper with the label used in NUMBER_TRACE.md), the manuscript source, and an MD5 manifest of every file. External inputs (the five human fixture datasets, SynthData 1.2.2 on npm, the RAITG harness at 10.5281/zenodo.20285103) are referenced with pinned versions and checksums rather than redistributed.

提供机构:
Zenodo
创建时间:
2026-09-24
二维码
社区交流群
二维码
科研交流群
商业服务