SocialIQA
收藏资源简介:
SocialIQA数据集由波恩大学对话式人工智能与社会分析实验室扩展,旨在评估语言模型在不同社会人口统计风格下的鲁棒性。该数据集包含1954个样本,源自SocialIQA验证集,涵盖了多种社会常识推理问题。数据集的创建过程通过LLAMA2模型生成不同人口统计风格的释义,确保语义相似度高于0.8。该数据集主要用于评估语言模型在复杂语言场景中的推理能力,特别是在面对不同人口统计风格的语言变化时的表现。
SocialIQA dataset was extended by the Conversational AI and Social Analysis Lab at the University of Bonn, aiming to evaluate the robustness of language models across different sociodemographic styles. This dataset includes 1,954 samples derived from the SocialIQA validation set, covering a variety of social commonsense reasoning questions. During the dataset construction, paraphrases in various demographic styles were generated using the LLAMA2 model, with their semantic similarity to the original texts maintained above 0.8. This dataset is primarily used to assess the reasoning capabilities of language models in complex linguistic scenarios, particularly their performance when facing language variations of different sociodemographic styles.




