ArGPT
收藏资源简介:
ArGPT是由圣保罗大学创建的一个新颖数据集,专注于分析ChatGPT生成的论点质量。该数据集包含168篇经过人工专家标注的辩论性文章,旨在通过模拟学生与教授的互动,生成并分析代表ChatGPT能力的论点。数据集内容涵盖多个领域,如艺术、历史、哲学和科学,每篇文章平均包含380个单词。创建过程中,通过精心设计的提示引导ChatGPT生成具有代表性的论点。ArGPT的应用领域主要集中在自动论点分类、论点挖掘和自动文章评分等任务,旨在解决大型语言模型在生成论点时可能出现的误导性问题,提供一个用于训练和测试相关系统的实用工具。
ArGPT is a novel dataset developed by the University of São Paulo, focusing on analyzing the quality of arguments generated by ChatGPT. This dataset contains 168 argumentative essays manually annotated by human experts, aiming to generate and analyze arguments representative of ChatGPT's capabilities by simulating interactions between students and professors. The dataset covers multiple domains such as art, history, philosophy and science, with each essay averaging 380 words. During its creation, ChatGPT was guided by carefully designed prompts to generate representative arguments. The main application fields of ArGPT focus on tasks including automatic argument classification, argument mining and automatic essay scoring, aiming to address the misleading issues that may occur when large language models generate arguments, and provide a practical tool for training and testing relevant systems.
ArGPT: 基于LLM的论证数据集
ArGPT数据集包含一组使用ChatGPT 3.5生成的论证性文章,并进行了以下标注:
- 论证挖掘:定义为三个不同的子任务,即跨度检测、组件分类和关系分类;
- 自动作文评分:使用真实世界论证性文章的修正标准;
- 论证质量:定义为好的文章如果用合理的论证来捍卫一个真实的声明,坏的文章如果论证有缺陷,或者丑的文章如果论证合理,但所支持的声明是错误的。评估标准包括:
- 标准0:明确陈述主要声明;
- 标准1:引入主题;
- 标准2:在文本中展开论证;
- 标准3:在结论中重述论证;
- 标准4:遵守标准语言规范;
- 标准5:正确使用论证连接词;
- 标准6:遵守主题;
- 标准7:不重复论证;
- 标准8:无矛盾;
- 标准9:不绕弯子;
- 标准10:陈述真实或合理的论证。
ArGPT v1:
包含168篇由单个标注者标注的文本。
ArGPT v2:
包含172篇文本,标注结果为两个不同标注者的共识。

- 1Assessing Good, Bad and Ugly Arguments Generated by ChatGPT: a New Dataset, its Methodology and Associated Tasks圣保罗大学 · 2024年



