Dataset for Deep Learning-Based Attribution of OpenAI GPT Models
收藏资源简介:
This dataset accompanies the paper "Deep Learning-Based Attribution of OpenAI GPT Models” and contains 100,000 text responses generated across four OpenAI large language models: GPT-3.5-Turbo, GPT-4-Turbo, GPT-4o, and GPT-4o-mini. Contents 25,000 unique prompts spanning 50 subject domains: Art & Culture, History, Nature, Science, and Technology 100,000 responses: 25k per model Each record includes structured fields: Category, Subcategory, Prompt variation, Prompt text, Prompt token count, Model name, Response text, Response token count, and Generation time Intended UseThe dataset was created to support research in authorship attribution, synthetic text analysis, and classifier benchmarking, and may be reused in broader studies of large language model behavior.



