JudgeGPT Human Perception Data: Dual-Axis Judgments of AI-Generated Disinformation
收藏资源简介:
Large-scale human perception data on AI-generated versus human-written text, collected through the JudgeGPT socio-technical evaluation platform. Participants assessed text fragments on two continuous axes, origin (human vs. machine) and veracity (legitimate vs. fake), while demographic, behavioral and temporal data were recorded. The deposit contains 2,546 judgments from 539 participants, together with the 3,278 stimulus fragments they were shown, so the judgments are self-contained and can be analyzed without a second download. Snapshot of 11 March 2026, unchanged. This version makes two corrections rather than adding data. First, the participant table is now a privacy-reduced publication copy. Four columns recorded during collection are withheld: IpLocation, UserAgent, ScreenResolution and QueryParams. Three derived columns replace them: IpCountry, IpContinent and RecruitmentRoute. Every figure the dissertation reports from this table remains reproducible. The earlier documentation stated that no directly identifying information was collected; that statement was wrong and has been corrected. Second, the bundled stimulus corpus is now byte-identical to RogueGPT Stimulus Corpus v1.2.0, which it previously was not. The copy shipped here carried 204 rows of double-encoded UTF-8, an outdated model identifier on two rows, one stray source URL on a machine-origin fragment, and an empty IngestedVia on 2,308 rows. Two deposits distributing the same corpus in two different states is a defect in its own right, and the two files are now checksum-identical by construction. Collected as part of a doctoral dissertation on the AI-driven disinformation ecosystem at Frankfurt University of Applied Sciences. Licensing and access This record is not covered by a single license, because the bundled stimulus corpus carries rights the depositor does not hold. The perception data (results.csv, participants-anonymized.csv) and all accompanying documentation are released under CC BY 4.0. Copyright 2024-2026 Alexander Loth. The bundled stimulus corpus (fragments.csv) has mixed rights, identical to the RogueGPT Stimulus Corpus deposit: the 2,638 machine-generated rows are CC BY 4.0, while the 640 human-sourced rows are excerpts of third-party news material whose copyright remains with the respective publishers. Those rows are made available for non-commercial academic text and data mining only, under the research exception of section 60d UrhG (Art. 3 EU DSM Directive). The Origin column separates the two parts. Participation was voluntary with informed consent. No names, email addresses or account identifiers were requested at any point, and ParticipantID is a random session identifier that resolves to nothing outside this dataset. No claim of anonymity is made for the withheld columns; they are withheld precisely because their combination was identifying.



