遇见数据集

SilentSpec Empirical Dataset: LLM-Generated TypeScript Unit Test Compile-Readiness Across Providers

收藏
Zenodo2026-04-12 更新2026-05-26 收录
官方服务:

资源简介:

Complete empirical dataset and analysis pipeline for: "Reliability of LLM-Generated Unit Tests: An Empirical Study of Compile-Readiness Across Providers and Frameworks" (Madduri, 2026). Contents:- 339 ground-truth records (288 T01 cold-start runs) evaluating compile-readiness of LLM-generated TypeScript unit tests across 3 providers (Claude claude-sonnet-4-6, OpenAI gpt-4o, GitHub Models gpt-4o), 43 source files, and 5 repositories (logic-bench, trpc, date-fns, prisma, typeorm).- Per-run metadata: provider, model, framework, sourceFile, repoName, compileReady, failureType, healerUsed, generationTimeMs, and 16 additional fields (schema v1.1).- Baseline comparison data: 120 raw LLM runs (no pipeline) across all 3 providers demonstrating 0-13.3% baseline CRR vs 62.5% pipeline CRR.- UZPR (User Zone Preservation Rate) validation: 30 runs with SHA-256 cryptographic hash verification confirming 100% byte-for-byte preservation of user-authored test content during active LLM-driven regeneration.- Statistical analysis results: Fisher's Exact Tests, Mann-Whitney U, Cohen's h effect sizes, Bonferroni-corrected p-values, Clopper-Pearson 95% confidence intervals.- Reproduction scripts: dataset aggregation, statistical analysis, baseline runner, compile-readiness checker.- 20 synthetic benchmark source files (logic-bench corpus).- 3 analysis charts (CRR by provider, CRR by repository, failure taxonomy).- Full README and step-by-step REPRODUCE.md. Associated tool: SilentSpec VS Code extension (Publisher: bharadwajmadduri.silent-spec)

提供机构:
Zenodo
创建时间:
2026-04-12
二维码
社区交流群
二维码
科研交流群
商业服务