How the Misuse of a Dataset Harmed Semantic Clone Detection
收藏资源简介:
This journal-specific supplementary package accompanies ‘How the Misuse of a Dataset Harmed Semantic Clone Detection’.BCB406 contains the 406 sampled clone pairs, snippets, investigation protocol and judgements, sampling summaries, uncertainty calculation, and five-run post-hoc LLM comparison with its majority table and offline calculation script.Literature contains the bibliography of 179 papers, historical prompts and request script, configuration documentation, and aggregate classification and agreement results. Individual evaluative classifications and identifiable literature LLM responses are retained internally and are available from the authors on request. Consequently, the literature agreement statistics cannot be independently recalculated from the public supplementary files alone.Embeddings contains 57 aggregate model–dataset results and their documentation. Pair-level scores and results using BigCloneBench-derived thresholds are not included. This is supplementary evidence, not a complete replication package.The shared GitHub repository separately retains the preliminary IWSC 2022 data. The IWSC2022 directory is intentionally excluded from this journal package. Earlier Git history and Zenodo versions remain unchanged as historical records.This revised version updates the sampling and uncertainty documentation, includes the two post-LLM clone-pair corrections, reports the final aggregate literature count of 128, and adds the aggregate embedding results.GitHub tag: https://github.com/jkrinke/BigCloneBench-Sample-Validation/tree/journal-v2Source commit: https://github.com/jkrinke/BigCloneBench-Sample-Validation/tree/04c2ea975c6e8f9660755935182ba518c45246e3



