BayLing-80
收藏资源简介:
BayLing-80数据集包含320条单轮和多轮指令,涵盖中文和英文。该数据集从将Vicuna评估中的80条英文指令翻译成中文开始,随后通过人工扩展生成了两种语言的单轮和多轮指令。该数据集主要用于评估大型语言模型(LLM)的跨语言和对话能力,覆盖了九项任务,包括写作、角色扮演、常识、费米问题、反事实问题、编程、数学、通用任务和知识。在评估过程中使用GPT-4进行打分。
The BayLing-80 dataset comprises 320 single-turn and multi-turn instructions in both Chinese and English. It was initially developed by translating 80 English instructions from the Vicuna evaluation benchmark into Chinese, followed by manual expansion to generate single-turn and multi-turn instructions for both languages. This dataset is primarily intended to evaluate the cross-lingual and conversational capabilities of Large Language Models (LLMs), covering nine task categories: writing, role-playing, common sense reasoning, Fermi problems, counterfactual questions, programming, mathematics, general tasks, and knowledge-related tasks. GPT-4 was utilized for scoring throughout the evaluation pipeline.




