HeshamHaroon/Arabic_fake_news_dataset
收藏资源简介:
Arabic_fake_news_dataset数据集是一个用于研究埃及社区中假新闻传播的新闻文章集合。该数据集包含从埃及平台[متصدقش (Matsda2sh)]抓取的新闻文章,分为假新闻和真新闻两类。数据集以JSON文件形式提供,每个新闻文章包含链接、假新闻标题列表和真新闻标题列表。数据集可能需要预处理步骤以确保数据质量和一致性,包括去除重复条目、处理缺失或错误数据、去除噪声或无关信息以及进行文本标记化和规范化。
The Arabic_fake_news_dataset is a collection of news articles intended for research on the spread of fake news in Egyptian communities. This dataset comprises news articles scraped from the Egyptian platform [متصدقش (Matsda2sh)], and is categorized into two classes: fake news and real news. It is provided in JSON file format, where each news article entry includes a link, a list of fake news headlines, and a list of real news headlines. Preprocessing steps may be required to ensure data quality and consistency, including removing duplicate entries, handling missing or erroneous data, eliminating noisy or irrelevant information, as well as performing text tokenization and normalization.
Arabic Fake News Dataset 概述
基本信息
- 语言: 阿拉伯语
- 名称: Arabic Fake News Dataset
- 标签:
- fake-news
- arabic
- web-scraping
- 任务类别:
- text-classification
- natural-language-processing
- web-scraping
- 许可证: Apache-2.0
数据集描述
- 来源: 数据集包含从埃及平台 متصدقش (Matsda2sh) 抓取的新闻文章。
- 目的: 用于研究和解决埃及社区中假新闻的传播问题。
- 内容: 包含被分类为假或真的新闻文章及其相应标题。
- 格式: 以 JSON 文件形式提供,文件名为
arabic_fake_news_dataset.json。 - 结构: 每个新闻文章以字典形式表示,包含
link(文章链接)、fakes(假新闻标题列表)和trues(真新闻标题列表)。
数据预处理
- 需求: 数据集需要进一步预处理以确保数据质量和一致性。
- 建议步骤:
- 移除重复条目。
- 处理缺失或错误数据。
- 去除网络抓取过程中引入的噪声或无关信息。
- 进行分词和文本规范化。
引用信息
-
作者: Hesham Haroon
-
年份: 2023
-
引用格式:
@misc{Arabic_fake_news_dataset, title = {Arabic_fake_news_dataset}, author = {Hesham Haroon}, year = {2023} }




