Extended Online Grooming Transcript Dataset
收藏资源简介:
This dataset is formed of three files: Convitions - found on the "peverted-justice.com" website under the convictions page, commonly used in literature as the PJ dataset. Archive - also found on the "peverted-justice.com" website however these appear to not be readily indexed as these pages were found through the discord channel. Discord - formed from collating .txt, .rtf, .doc/x, and .zip archives. These files were made accessible through a "discord server" (https://discord.gg/9bZFVTPHRh) which was created by individuals who work with "peverted-justice.com". These .csv files are intended to be processed as dataframes using the pandas package in python (pandas - Python Data Analysis Library), and read using the "read_csv" function. These files follow the following format: Transcript Num Message Num Actor Message Line A unique identifier for each transcript Denotes the position of the Message Line within the Transcript Either 1 (Attacker/Predator) or 0 (Child/Honeypot) A pre-processed utterance (message)



