GPT Reddit Dataset (GRiD)
收藏资源简介:
GPT Reddit Dataset (GRiD) 是由加州大学河滨分校创建的一个用于检测GPT生成文本的数据集。该数据集包含6513个样本,其中1368个由GPT-3.5-turbo模型生成,5145个由人类生成。数据来源于Reddit和OpenAI API,通过特定的收集和处理流程确保数据的质量和区分度。GRiD旨在为评估和提升GPT文本检测技术提供基准,解决互联网上AI驱动通信的信任和责任问题。
GPT Reddit Dataset (GRiD) is a dataset dedicated to detecting GPT-generated texts, developed by the University of California, Riverside. It consists of 6,513 samples, with 1,368 generated by the GPT-3.5-turbo model and 5,145 generated by human users. The data is sourced from Reddit and the OpenAI API, and its quality and discriminative performance are ensured through a targeted collection and processing workflow. GRiD aims to serve as a benchmark for evaluating and advancing GPT text detection technologies, and to address the trust and accountability issues surrounding AI-driven communications on the Internet.

- 1GPT-generated Text Detection: Benchmark Dataset and Tensor-based Detection Method加州大学河滨分校 · 2024年



