Bias Benchmark for Generation (BBG)
收藏资源简介:
Bias Benchmark for Generation (BBG)是一个用于评估大型语言模型(LLM)在社会偏见方面的基准数据集,由KAIST的研究人员构建。该数据集基于英语和韩语的BBQ(Bias Benchmark for QA)数据集,通过替换故事情境中的人物描述为中性的占位符,来评估LLM在长篇故事生成中的偏见。BBG包含9个类别的232个模板和12个类别的286个模板,分别对应英语和韩语版本,共计120508个故事和问题对。该数据集旨在解决LLM在长篇生成中的社会偏见评估问题,推动公平的自然语言处理系统的发展。
Bias Benchmark for Generation (BBG) is a benchmark dataset for evaluating social biases in Large Language Models (LLMs), constructed by researchers from KAIST. This dataset is built upon the English and Korean versions of the BBQ (Bias Benchmark for QA) dataset, and assesses biases in LLMs' long-form story generation by replacing character descriptions in story contexts with neutral placeholders. BBG contains 232 templates across 9 categories and 286 templates across 12 categories, corresponding to the English and Korean versions respectively, with a total of 120,508 story-question pairs. This dataset aims to address the issue of social bias evaluation in LLMs' long-form generation, and promote the development of fair natural language processing systems.

- 1Social Bias Benchmark for Generation: A Comparison of Generation and QA-Based EvaluationsKAIST · 2025年



