Wikidata Vandalism Detection Dataset
收藏资源简介:
Description This dataset accompanies a research paper that introduces a new system designed to support the Wikidata community in combating vandalism on the platform. Keywords: Wikidata, content differences, vandalism detection, data mining, content analysis, computational social science, NLP. Dataset Details: Number of files: 20 (6.85 GB) Format: CSV License: CC BY 4.0 Use Case: data mining, vandalism detection and analysis, content moderation. The dataset is primarily intended for training and evaluating vandalism detection systems for Wikidata. Observation period: 21 months of training, 3 months hold-out testing. Data from 01.09.2021 to 01.09.2023 (Snapshot from 2024-04) Features: Each record characterizes the corresponding revision of the Wikidata record, including revision metadata, user details, content modifications (insert, remove, or change), and corresponding MLMs-based features. Data Filtering and Feature Engineering: Advanced filtering and feature engineering techniques were applied to ensure the dataset's quality and relevance for effectively training the vandalism detection system. Files: 2024-04_content_batch_{i}.csv - Revision content features (split into 15 batches). 2024-04_metadata.csv - Revision metadata features. expert_scores.csv - revision labeled by expert (column label correspond to the expert label). full_labels_2024-04_text_en.csv - Wikidata ID to English label mapping. mlm_text_features.csv - pretrained MLM scores. ores_scores.csv - ORES (previous model in production) scores. Attribution The dataset was compiled from the Wikidata dump. All structured data from the Wikidata main, Property, Lexeme, and EntitySchema namespaces is available under the Creative Commons CC0 License; text in the other namespaces is available under the Creative Commons Attribution-ShareAlike License. Related paper citation: TBD



