遇见数据集

Replication package for "Instruments Before Results: A Pre-Registered Study of LLM Refinement of Decompiled Code"

收藏
Zenodo2026-08-16 更新2026-08-20 收录
官方服务:

资源简介:

A pre-registered study of LLM refinement of decompiler output: five transferable instrument findings, one bounded result, no confirmed registered hypothesis. The result is a memorization floor, measured within-item so corpus difficulty cannot contribute: on functions written after the analysis plan was committed, identifier recovery survives destruction of the input's dataflow (95% CI [-0.026, +0.026], n=12). Recovery is real but modest -- +0.072 above an arm-matched permutation null -- and the ablation leaves operations intact, so naming from operations alone remains a competing reading. At full corpora all six contrasts are null, and a second refiner from another vendor -- registered in advance, on byte-identical inputs -- reproduces this on wider bounds: twelve contrasts, two models, twelve nulls. The findings:(i) harness bias -- three bugs in our reassembly path manufactured a 20-point correctness deficit, each penalizing the refined arm for obeying the prompt; (ii) judge divergence -- our LLM readability judge ranks arms as a hypothesis-blind human rater does but reports a 2.8x larger gap, concentrated in the raw baseline; (iii) non-applicability -- our registered equivalence checker did not discriminate on optimized output, leaving three hypotheses untestable; (iv) execution-specific precision -- a registered re-execution reproduced every verdict while widening the null's bound 44%; and (v) null mis-specification -- our chance baseline paired ground truth with ground truth, making realsignal read as chance. Our ablation as first implemented did not ablate. We recommend the floor, manipulation checks, arm-matched baselines, applicability reports, judge calibration and registered re-execution.

提供机构:
Zenodo
创建时间:
2026-08-16
二维码
社区交流群
二维码
科研交流群
商业服务