SOCIALHARMBENCH
收藏资源简介:
SOCIALHARMBENCH是一个包含585个提示的数据集,跨越7个社会政治类别和34个国家,旨在揭示大型语言模型(LLMs)在政治敏感环境中最易出现的问题。该数据集由多伦多大学、向量研究所等多家机构的研究人员共同创建,旨在评估LLMs在处理社会政治危害方面的安全性。数据集内容涵盖了从19世纪至今的历史事件,并跨越了所有大洲的34个国家,包括德国、美国、中国、俄罗斯/苏联和柬埔寨等国家。数据集的创建过程包括确定历史事件、策划子主题、将子主题与历史事件相结合以及生成模板,以确保评估能够反映社会恶意,而不是语义影响。SOCIALHARMBENCH的目标是评估LLMs在社会政治危害方面的安全性,旨在解决当前安全措施在高风险社会政治环境中的不足。
SOCIALHARMBENCH is a dataset containing 585 prompts, spanning 7 socio-political categories and covering 34 countries across all continents. It is designed to uncover the most prevalent vulnerabilities of large language models (LLMs) in politically sensitive contexts. This dataset was co-developed by researchers from multiple institutions including the University of Toronto and the Vector Institute, with the core goal of evaluating the safety of LLMs when addressing socio-political harm. The dataset covers historical events from the 19th century to the present, and includes countries such as Germany, the United States, China, Russia/the Soviet Union, and Cambodia, among others. The construction process of SOCIALHARMBENCH includes identifying historical events, curating sub-themes, combining sub-themes with historical events, and generating evaluation templates, to ensure that the assessment reflects social maliciousness rather than mere semantic impacts. The primary objective of SOCIALHARMBENCH is to assess the safety of LLMs against socio-political harm, addressing the shortcomings of current safety measures in high-stakes socio-political scenarios.




