WORLDVALUESBENCH
收藏资源简介:
WORLDVALUESBENCH是一个全球多样性的大规模基准数据集,用于多文化价值感知语言模型的研究。该数据集源自世界价值观调查(WVS),涵盖了来自全球94,728名参与者的数百个价值观问题的答案。数据集包含超过2000万个“(人口统计属性, 价值问题) → 答案”类型的示例。创建过程涉及从WVS响应中提取和构造数据,重点关注不同价值观问题和详细的人口统计属性。该数据集的应用领域主要集中在研究语言模型在多文化价值预测任务中的局限性和机会,旨在提升语言模型在生成安全和个性化响应方面的能力。
WORLDVALUESBENCH is a large-scale benchmark dataset with global diversity, designed for research on multicultural value-aware language models. The dataset is derived from the World Values Survey (WVS), covering answers to hundreds of value-related questions from 94,728 participants across the globe. The dataset contains over 20 million examples of the format "(demographic attribute, value question) → answer". Its construction process involves extracting and structuring data from WVS responses, with a focus on diverse value-related questions and detailed demographic attributes. Its main application scenarios focus on investigating the limitations and opportunities of language models in multicultural value prediction tasks, aiming to enhance the ability of language models to generate safe and personalized responses.

- 1WorldValuesBench: A Large-Scale Benchmark Dataset for Multi-Cultural Value Awareness of Language Models马萨诸塞大学阿默斯特分校 · 2024年



