Nunchi-Bench
收藏资源简介:
Nunchi-Bench 是一个用于评估大型语言模型 (LLMs) 文化理解和推理能力的基准数据集,专注于韩国的迷信。数据集包含 247 个问题,涵盖了 31 个主题,评估了模型对韩国迷信的事实知识、文化适宜性建议和情境解释能力。数据集包括三种类型的任务:多项选择题 (MCQs) 评估对韩国迷信的事实知识;陷阱问题评估模型在文化敏感情境中提供适宜建议的能力;解释问题检验模型是否能够从社交互动中推断文化意义。Nunchi-Bench 同时提供韩文和英文版本,以促进多语言模型的评估。此外,对于陷阱和解释任务,还提供了明确指定或省略韩国文化背景的版本。该数据集旨在帮助研究人员评估和改进 LLMs 在跨文化环境中的表现,特别是在处理文化情境时。
Nunchi-Bench is a benchmark dataset designed to evaluate the cultural understanding and reasoning capabilities of Large Language Models (LLMs), with a particular focus on Korean superstitions. The dataset comprises 247 questions spanning 31 distinct topics, and assesses models' factual knowledge of Korean superstitions, capacity to provide culturally appropriate advice, and skills in contextual interpretation. The dataset includes three task categories: Multiple Choice Questions (MCQs) for evaluating factual knowledge of Korean superstitions; trap questions for testing the model's ability to deliver suitable advice in culturally sensitive scenarios; and explanation questions that examine whether the model can infer cultural meanings from social interactions. Nunchi-Bench is offered in both Korean and English versions to support the evaluation of multilingual models. Furthermore, versions with explicitly specified or omitted Korean cultural contexts are provided for the trap and explanation tasks. This dataset is intended to help researchers evaluate and improve the performance of LLMs in cross-cultural settings, particularly when engaging with cultural contexts.
Nunchi-Bench数据集概述
数据集基本信息
- 数据集名称:Nunchi-Bench
- 托管平台:GitHub
数据集描述
(注:根据提供的README内容,该数据集详情页未包含具体描述信息)

- 1Nunchi-Bench: Benchmarking Language Models on Cultural Reasoning with a Focus on Korean Superstition洛桑联邦理工学院 (EPFL) 和 首尔国立大学 (SNU) · 2025年



