Parallel Communities Across the Surface Web and the Dark Web
收藏资源简介:
DATA DESCRIPTION This dataset supports a comparative analysis of online communities across two distinct platforms: Reddit (a mainstream forum) and Dread (a dark web discussion forum). It includes: 7M+ posts and comments from Reddit 200K+ posts and comments from Dread The data is pre-processed and aligned across similar topics. We provide hashed user identifiers, toxicity scores (via Detoxify), timestamps (in UTC), and cleaned text content. Personally identifiable information has been removed or anonymized. Field Name Field Description id A unique identifier for each entry in the dataset. parent_id Identifier used when the entry is part of a conversation thread or linked to a comment; used to associate replies with their parent. processed_text The text content of the comment or post, pre-processed (e.g., cleaned, normalized) for analysis. score / vote The numeric score or vote count associated with the entry. timestamp The UTC timestamp indicating when the entry was created. subreddit_name / community_name The name of the community (e.g., subreddit) where the entry was posted. hashed_user_id A pseudonymized identifier for the user who created the entry, generated using a salted hash. No original usernames are retained. toxic A numerical score indicating the level of toxicity (https://huggingface.co/unitary/toxic-bert) DATA ACCESS INTRUCTIONS This dataset is under restricted access. To request access, please follow the steps below: Login to Zenodo account. On the Files section, enter your details (email and name). In the message box state your institutional affiliation and provide a brief description of your intended use of the data. Click on "Request accces". Your request will be reviewed within 3–5 business days. You will receive an email notification once access has been granted. If you have questions or need support with your request, please contact us at:megha[dot]sundriyal[at]mpi-sp[dot]org ETHICAL USAGE COMMITMENT By requesting access, users acknowledge and agree to use the dataset solely for ethical and lawful research purposes. The following uses are strictly prohibited: Developing tools or algorithms intended to promote or assist in illicit activities or to evade law enforcement or regulatory oversight. Any form of algorithmic discrimination or bias based on protected characteristics such as race, gender, sexual orientation, religion, or political beliefs. Unauthorized access, data leakage, or any action that compromises the confidentiality or integrity of the dataset.



