TGEA
收藏资源简介:
TGEA数据集是由天津大学和华为诺亚方舟实验室共同构建的,首个基于机器生成的文本的错误注释数据集,包含多个基准任务,用于评估预训练语言模型在文本生成方面的能力。该数据集从中文GPT-2模型生成的文本中收集原始数据,通过人工标注,检测出错误句子,并进一步进行错误类型、相关文本跨度、错误纠正以及错误原因的标注。数据集涵盖24种错误类型的双层错误分类体系,旨在促进对预训练语言模型生成的文本进行自动错误检测和纠正的研究。
The TGEA dataset, co-developed by Tianjin University and Huawei Noah's Ark Lab, is the first error-annotated dataset for machine-generated text. It includes multiple benchmark tasks designed to evaluate the text generation capabilities of pre-trained language models. Raw data for the dataset is collected from texts generated by Chinese GPT-2 models, where erroneous sentences are identified via manual annotation, followed by further annotations of error types, relevant text spans, error corrections, and error causes. The dataset features a two-tier error classification system covering 24 error types, and aims to advance research on automatic error detection and correction for texts generated by pre-trained language models.




