MORAL INTEGRITY CORPUS (MIC)
收藏资源简介:
MORAL INTEGRITY CORPUS(MIC)是由乔治亚理工学院创建的一个大型数据集,包含38,000个对话系统回复的人类编写提示。该数据集通过99,000个不同的经验法则(RoTs)捕捉了对话系统的道德假设,每个RoT反映了特定的道德信念,解释了聊天机器人的回复为何可能被视为可接受或有问题。数据集的创建过程涉及从大量公开数据中筛选和标注,确保内容的质量和相关性。MIC数据集主要用于理解和基准测试对话代理在开放领域“闲聊”设置中的隐含道德假设和灵活的常识推理能力,旨在解决对话系统中的道德和社交常识推理问题。
MORAL INTEGRITY CORPUS (MIC) is a large-scale dataset developed by the Georgia Institute of Technology, encompassing 38,000 human-written prompts for dialogue system responses. This corpus captures the moral assumptions inherent in dialogue systems via 99,000 distinct Rules of Thumb (RoTs), with each RoT representing a specific moral belief and explaining why a chatbot’s response could be deemed acceptable or problematic. The development process of the MIC dataset entailed screening and annotating data from a large pool of public sources to guarantee its content quality and relevance. Primarily, the MIC dataset is utilized to understand and benchmark the implicit moral assumptions and flexible commonsense reasoning abilities of dialogue agents within open-domain "small talk" scenarios, with the goal of resolving challenges concerning moral and social commonsense reasoning in dialogue systems.

- 1The Moral Integrity Corpus: A Benchmark for Ethical Dialogue Systems乔治亚理工学院 · 2022年



