BAEM1N/Korean-RAG-LLM-Judge-Benchmark
收藏资源简介:
该数据集是韩国RAG(检索增强生成)LLM作为法官交叉验证基准,旨在通过46个LLM和9个法官的交叉验证矩阵评估韩国RAG答案的质量。数据集基于allganize/RAG-Evaluation-Dataset-KO构建,增加了受控检索、交叉验证分数和RRF融合排名。数据集包含多个部分(问答、检索、候选答案、法官评分),并提供了每部分的详细列说明。此外,还描述了方法学、文件结构、使用示例以及许可和引用信息。
This dataset is the Korean RAG LLM-as-Judge Cross-Validation Benchmark, designed to evaluate the quality of Korean RAG (Retrieval-Augmented Generation) answers through a cross-validation matrix involving 46 LLMs and 9 judges. The dataset is built upon the allganize/RAG-Evaluation-Dataset-KO, adding controlled retrieval, cross-validation scores, and RRF-fused rankings. It includes various splits (qa, retrieval, cand_answers, judge_scores) and provides detailed column specifications for each. The methodology, file structure, usage examples, licensing, and citation information are also described.



