thai-bar-exam-judging
收藏资源简介:
Thai Bar-Exam Judging Corpus是一个泰语法律论述数据集,源自律师考试准备练习,旨在支持大型语言模型(LLM)作为评分者与人类考官之间稳定性的对比研究。数据集包含匿名化的自由形式泰语法律论述,每篇论述由三位经过泰国律师理事会培训的考官进行评分,并对约三分之二的答案提供了基于跨度的内联评论。八名LLM考生在相同条件下参加相同考试,其答案由相同考官进行盲评。此外,150个答案中的15个由两位非主要考官进行交叉评分,形成了一个三评分者稳定性子集,用于相关论文分析。数据规模约为2.3 MB,包含问题文本、人类考生论述、LLM考生论述、交叉评分和基于跨度的评论等配置文件,以及宽格式三评分者分数矩阵和预计算的人类评分者间一致性指标等衍生文件。数据集适用于泰语法律NLP基准测试、基于跨度的评论建模、评估者间一致性研究以及文本分类和生成任务,但规模较小(基于三个商法问题),不适用于大规模部署评分或分类器训练。数据经过匿名化处理以保护隐私,基于CC-BY 4.0许可证发布。
The Thai Bar-Exam Judging Corpus is a Thai legal argumentation dataset derived from bar exam preparation exercises, aimed at supporting comparative studies on the stability between large language models (LLMs) as scorers and human examiners. It contains anonymized free-form Thai legal arguments, each scored by three examiners trained by the Thai Lawyers Council, with inline span-based comments provided for approximately two-thirds of the answers. Eight LLM candidates took the same exam under identical conditions, and their answers were blindly scored by the same examiners. Additionally, 15 out of 150 answers were cross-scored by two non-primary examiners, forming a three-scorer stability subset for related paper analysis. The dataset has a total size of approximately 2.3 MB, including configuration files such as question texts (3 lines), human candidate arguments (126 lines, corresponding to 42 human candidates × 3 questions), LLM candidate arguments (24 lines, corresponding to 8 LLM models × 3 questions), cross-scoring (30 lines, corresponding to 15 cross-scoring units × 2 cross-scorers), and span-based comments (1164 lines, including 1079 primary comments and 85 cross-comments). It also provides derived files, such as wide-format three-scorer score matrices and precomputed inter-rater agreement metrics (Krippendorff alpha). The dataset is suitable for various tasks, including Thai legal NLP benchmarking, span-based comment modeling, inter-rater agreement studies, and text classification and generation tasks. The data is anonymized to protect privacy, with both scorers and candidates hidden, and is released under the CC-BY 4.0 license. It is applicable for stability research methodology but, due to its small scale (based on three commercial law questions), is not suitable for large-scale deployment of scoring or classifier training.




