Goofus & Gallant Story Corpus
收藏资源简介:
Goofus & Gallant故事语料库是一个多模态数据集,旨在通过自然语言文本和图像展示社会价值观。该数据集由肯塔基大学和佐治亚理工学院的研究团队创建,基于儿童漫画《Goofus & Gallant》,该漫画通过对比两种行为模式(规范与非规范)来教育儿童社会原则。数据集包含1387条文本和819张图像,涵盖了1995年至2017年的漫画内容。数据集的创建过程包括从漫画中提取文本和图像,并通过众包和大型语言模型(LLMs)对行为进行社会原则的标注。该数据集主要用于训练AI系统,使其能够理解和遵循人类社会的价值观,解决AI与人类价值观对齐的问题。
The Goofus & Gallant Story Corpus is a multimodal dataset designed to showcase social values through natural language text and images. This dataset was developed by research teams from the University of Kentucky and the Georgia Institute of Technology, based on the children's comic strip *Goofus & Gallant*, which educates children about social principles by contrasting two sets of behavioral patterns: normative and non-normative ones. The dataset comprises 1,387 text entries and 819 images, spanning comic strip content from 1995 to 2017. The process of constructing the dataset includes extracting text and images from the comic strips, and annotating the behaviors with social principles via crowdsourcing and Large Language Models (LLMs). This dataset is primarily intended for training AI systems to understand and comply with human social values, addressing the challenge of AI alignment with human values.




