遇见数据集

BAEM1N/Korean-RAG-LLM-Judge-Benchmark

收藏
Hugging Face2026-04-28 更新2026-05-03 收录
官方服务:

资源简介:

该数据集是韩国RAG(检索增强生成)LLM作为法官交叉验证基准,旨在通过46个LLM和9个法官的交叉验证矩阵评估韩国RAG答案的质量。数据集基于allganize/RAG-Evaluation-Dataset-KO构建,增加了受控检索、交叉验证分数和RRF融合排名。数据集包含多个部分(问答、检索、候选答案、法官评分),并提供了每部分的详细列说明。此外,还描述了方法学、文件结构、使用示例以及许可和引用信息。

This dataset is the Korean RAG LLM-as-Judge Cross-Validation Benchmark, designed to evaluate the quality of Korean RAG (Retrieval-Augmented Generation) answers through a cross-validation matrix involving 46 LLMs and 9 judges. The dataset is built upon the allganize/RAG-Evaluation-Dataset-KO, adding controlled retrieval, cross-validation scores, and RRF-fused rankings. It includes various splits (qa, retrieval, cand_answers, judge_scores) and provides detailed column specifications for each. The methodology, file structure, usage examples, licensing, and citation information are also described.

提供机构:
BAEM1N
二维码
社区交流群
二维码
科研交流群
商业服务