GREEKBARBENCH
收藏资源简介:
GREEKBARBENCH是一个针对法律推理和引用的基准数据集,包含来自希腊律师考试的五个不同法律领域的自由文本问题。数据集要求引用法律条文和案件事实。为了解决自由文本评估的挑战,我们提出了一个三维评分系统,并结合了LLM-as-a-judge的方法。我们还开发了一个元评估基准,以评估LLM-judges与人类专家评估之间的相关性,结果表明,简单的基于跨度的评分标准提高了它们的对齐度。我们对13个专有和开放权重LLMs的系统评估表明,尽管最好的模型超过了平均专家分数,但它们仍然低于专家分数的第95百分位数。
GREEKBARBENCH is a benchmark dataset focused on legal reasoning and citation, consisting of free-text questions across five distinct legal domains sourced from the Greek bar examination. The dataset requires respondents to cite legal statutes and case facts. To address the challenges of free-text evaluation, we propose a three-dimensional scoring system combined with the LLM-as-a-judge methodology. We additionally develop a meta-evaluation benchmark to assess the correlation between LLM-judges and human expert evaluations, with results demonstrating that simple span-based scoring criteria improve their alignment. Our systematic evaluation of 13 proprietary and open-weight LLMs shows that while the top-performing models exceed the average expert score, they still fall below the 95th percentile of expert scores.




