A Diagnostic Framework for LLMs as Code Judges: Understanding Misjudgments
收藏官方服务:
资源简介:
This repository contains the scripts used to build and analyze a misjudgement-focused version of the CodeJudge-Eval benchmark. Our workflow augments CodeJudge-Eval dataset with static code quality metrics and problem level metrics. Additionally, the repository contains the code and the data used for the analysis done in the study to answer the following research questions: RQ1: Taxonomy of Misjudgments RQ2: Predictability of LLM Misjudgments from Code-level and Problem-level Features RQ3: Understanding Feature Influence using SHAP-Based Explanations
提供机构:
Zenodo创建时间:
2026-03-24



