Arabic-news-daily
收藏资源简介:
Arabic News Daily 是一个每日自动更新的阿拉伯语新闻数据集,从 Al Jazeera Arabic、BBC Arabic、RT Arabic 和 Al Arabiya 等主要阿拉伯语新闻来源收集。与其他静态快照数据集不同,该数据集每天增长,持续提供最新鲜的阿拉伯语文本,适用于需要实时语料的研究。每条记录包含以下字段:唯一标识符(id,基于URL的MD5哈希)、标题(title)、清洗后的正文(content)、原始URL(url)、来源名称(source,包括 aljazeera、bbc_arabic、rt_arabic、alarabiya)、发布时间(published_at,RFC 2822格式)和采集时间(collected_at,ISO 8601 UTC时间戳)。数据集通过自动化流水线在每天UTC 03:00更新,规模随时间增长,当前约有一万至十万条样本。该数据集可用于阿拉伯语语言模型预训练/微调、阿拉伯语NLP基准测试、新闻摘要、命名实体识别(阿拉伯语)以及主题分类等任务。采用CC BY 4.0许可证。
Arabic News Daily is a daily automatically updated Arabic news dataset collected from major Arabic news sources such as Al Jazeera Arabic, BBC Arabic, RT Arabic, and Al Arabiya. Unlike other static snapshot datasets, this dataset grows daily, continuously providing the freshest Arabic text, suitable for research requiring real-time corpora. Each record contains the following fields: a unique identifier (id, MD5 hash based on URL), title, cleaned text content, original URL, source name (including aljazeera, bbc_arabic, rt_arabic, alarabiya), publication time (published_at, RFC 2822 format), and collection time (collected_at, ISO 8601 UTC timestamp). The dataset is updated via an automated pipeline daily at 03:00 UTC, with its size growing over time, currently around 10,000 to 100,000 samples. This dataset can be used for Arabic language model pre-training/fine-tuning, Arabic NLP benchmarks, news summarization, named entity recognition (Arabic), and topic classification. It is licensed under CC BY 4.0.
数据集名称:Arabic News Daily 🗞️
数据集简介:这是一个每日更新的阿拉伯语新闻数据集,由系统自动从多家主要阿拉伯语新闻源收集。与静态快照型数据集不同,该数据集会每日增长,适合需要最新阿拉伯语文本的研究工作。
语言与许可证
- 语言:阿拉伯语(ar)
- 许可证:CC BY 4.0
数据规模与划分
- 样本数量:12.2万条(训练集)
- 数据大小:约140 KB(数据下载大小约63 KB)
- 规模类别:10K < n < 100K
- 数据划分:仅包含训练集(train),共122条样本
数据来源 数据集从以下四个主要阿拉伯语新闻源收集:
- Al Jazeera Arabic(现代标准阿拉伯语)
- BBC Arabic(现代标准阿拉伯语)
- RT Arabic(现代标准阿拉伯语)
- Al Arabiya(现代标准阿拉伯语)
数据集结构 每条记录包含以下字段:
- id:URL的MD5哈希值
- title:文章标题
- content:文章正文(经过清洗处理)
- url:文章链接
- source:来源标识(aljazeera | bbc_arabic | rt_arabic | alarabiya)
- published_at:发布日期(RFC 2822格式)
- collected_at:收集时间(ISO 8601 UTC时间戳)
更新机制 数据集通过自动化流水线(运行于自托管家庭实验室)于每天UTC时间03:00进行收集和推送更新。
适用任务 该数据集适用于以下NLP任务:
- 阿拉伯语语言模型预训练/微调
- 阿拉伯语NLP基准测试
- 新闻摘要研究
- 阿拉伯语命名实体识别
- 主题分类
引用方式 如需引用该数据集,请使用以下BibTeX格式:
bibtex @dataset{unohamza_arabic_news_daily, author = {unohamza}, title = {Arabic News Daily}, year = {2026}, publisher = {Hugging Face}, url = {https://huggingface.co/datasets/unohamza/Arabic-news-daily} }




