TBML-NoiseBench: Enforcement-Anchored Gold Triads, Dosage-Graded Episode Grid, and CLEA Evaluation Code
收藏资源简介:
TBML-NoiseBench is an enforcement-anchored benchmark for measuring how retrieval-augmented generation (RAG) over trade-based money-laundering (TBML) evidence degrades under adversarially contaminated retrieval. The release contains 184 gold triads across ten TBML typologies, a dosage-graded grid of frozen retrieval episodes that varies the ratio of hard, style-matched adversarial spans to genuine evidence, and CLEA, a claim-level evidence-attribution metric scored against human-authored gold rather than against retrieved context. It accompanies the paper of the same name under review at the Machine Learning journal via the ACML 2026 Journal Track. Version 1.0.2 (2 September 2026). Adds the full-panel judge-swap at the diagnostic dose. Nothing from version 1.0.1 is removed or altered. Added:- results/panel_results_swap_extension.jsonl: every model's diagnostic-dose (d3_high) responses on the interleave frame re-scored by the alternate-family judge, 2,154 rows (one carries an error and no scores). With the swap rows already in panel_results.jsonl this gives 2,569 doubly-judged responses across all fourteen models. Produced by code/extend_judge_swap.py on 2 September 2026 using the judge prompt, parser and CLEA function copied verbatim from code/run_panel_x.py; judges gpt-4o and claude-sonnet-4-5-20250929.- code/extend_judge_swap.py: the script above. Its --analyze mode prints two frames: the three-model scope of the original submission (reproducing its 0.020, 0.003 and 0.010 and alpha 0.792 on 1,246 responses) and the full-panel interleave frame (alpha 0.817; per-model differences at most 0.056, above 0.02 for four models; Spearman 0.79 between judge orderings).- docs/TBML-NoiseBench_paper.pdf and docs/TBML-NoiseBench_supplementary.pdf updated to the versions that report the full-panel swap.- docs/CHANGELOG_v1.0.2.md. Correction to the version 1.0.0 and 1.0.1 manuscripts: the judge effect on CLEA-F was stated as bounded at or below 0.02 on the three-model swap. On the full panel the per-model difference reaches 0.056; the manuscript now reports the measured values.



