遇见数据集

JoNewsFake dataset

收藏
Zenodo2026-03-26 更新2026-05-26 收录
官方服务:

资源简介:

JoNewsFake is a large-scale, multi-level Arabic news dataset curated from the official Facebook pages of 12 verified Jordanian news agencies. The dataset was collected, filtered, pre-processed, and annotated as part of a doctoral research project focused on multi-level Arabic fake news detection. It contains 50,000 posts after cleaning and filtering, representing one of the most comprehensive Arabic news datasets available for misinformation research. The dataset provides three layers of annotation: Main Category (22 classes)Broad thematic categories such as Politics, Economy, Health, Education, Culture, Government, etc. Sub-Category (75+ classes)Fine-grained hierarchical labels capturing the specific domain of each news post. Fake / Real Label (Binary)Posts are classified as Fake or Real based on a structured multi-stage annotation workflow. Each news post includes: Cleaned Arabic text (after normalization, noise removal, and preprocessing) Part-of-Speech (POS) tags Emotional indicators (26 features) AraBERT-based semantic embeddings (768-dim) — not included directly due to size, but can be reproduced using provided scripts Metadata fields such as post ID, source agency, and timestamp (when available) 🟦 Data Collection and Annotation Data were collected using the Facepager tool from 12 major Jordanian news outlets.The initial 134,160 posts were filtered according to: Minimum text length Removal of non-Arabic content Removal of duplicated or low-information posts Exclusion of advertisements or irrelevant material A rigorous six-person annotation protocol was used: Four trained annotators labeled the data. Two domain experts validated the labels. Disagreements were resolved through consensus. Inter-annotator reliability was measured during pilot rounds. 🟦 Pre-processing Pipeline The dataset underwent an extensive Arabic-specific pre-processing stage, including: Normalization (Hamza, Alef forms, Ta Marbuta → Ha) Tokenization Removal of diacritics, emojis, URLs, and noise POS tagging using Farasa / StanfordNLP Generation of emotional indicators using dictionary-based mapping This hybrid representation (text + POS + emotions) enables both classical ML and deep learning applications. 🟦 Intended Use This dataset is designed for research in: Arabic fake news detection Multi-label and multi-level text classification Hierarchical NLP modeling Representation learning for Arabic Evaluation of machine learning, deep learning, and transformer-based models It has already been used in experiments comparing ensemble models (Random Forest, Extra Trees, XGBoost, LightGBM) and deep learning models (CNN-BiLSTM, CNN-BiGRU, Transformer Encoder).

提供机构:
Zenodo
创建时间:
2026-03-26
二维码
社区交流群
二维码
科研交流群
商业服务