just-eval-instruct-eu
收藏资源简介:
Just-Eval EUS是一个用于指令响应质量评估的巴斯克语基准数据集。它是`re-align/just-eval-instruct`数据集的巴斯克语适配版本,专门为研究低资源语言环境下的大语言模型自动评估能力而构建。该数据集源自论文《Judging Instruction Responses in a Low-Resource Language: A Case Study on Basque》,通过对原英文Just-Eval基准进行翻译和人工后编辑得到。数据以JSON Lines格式存储,样本数量在1,000到10,000条之间。核心任务是评估模型对给定指令所生成回答的质量,关注回答的细粒度属性,并扩展了评判回答语言一致性和语法正确性的评估维度。该数据集旨在服务于研究在巴斯克语这类低资源场景下,不同大语言模型作为自动评判者的性能表现、它们与人类评判结果的相关性,以及人类评判本身作为可靠基准的可行性评估。
Just-Eval EUS is a Basque benchmark dataset for instruction response quality evaluation. It is a Basque-adapted version of the `re-align/just-eval-instruct` dataset, specifically constructed for investigating the automatic evaluation capabilities of large language models (LLMs) in low-resource language scenarios. This dataset originates from the paper *Judging Instruction Responses in a Low-Resource Language: A Case Study on Basque*, and was developed by translating the original English Just-Eval benchmark and performing manual post-editing. The data is stored in JSON Lines format, with the number of samples ranging from 1,000 to 10,000. Its core task is to evaluate the quality of responses generated by models for given instructions, focusing on the fine-grained attributes of the responses, and expands the evaluation dimensions for judging the linguistic coherence and grammatical correctness of the responses. This dataset aims to support research on the performance of different large language models as automatic evaluators in low-resource settings like Basque, their correlation with human evaluator results, and the feasibility assessment of human evaluation itself as a reliable benchmark.





