遇见数据集

AraMuST files before normalization old version

收藏
Zenodo2026-03-07 更新2026-05-26 收录
官方服务:

资源简介:

AraMuST is a large-scale unified Arabic dataset designed for multi-task learning across text summarization and topic classification tasks.It consolidates and harmonizes three major publicly available resources XL-Sum, SANAD, and NADCG under a single standardized schema. The dataset provides high-quality text–summary pairs and consistent topic labels across 8 major categories (Sports, Politics, Finance, Technology, Medical, Culture, Religion, and Accidents). Rigorous quality filtering and multi-step validation were conducted to ensure reliability and prevent data leakage. Duplicates across datasets were detected and removed using MD5 hashing and embedding-based cosine similarity. Each entry in AraMuST follows a unified JSONL structure:{"task": "classification|summarization|classification+summarization", "text": str, "summary": str|NA, "label": str|NA, "source": str} Composition:Total samples: 2,143,833Tasks: Summarization, Classification, and Joint Summarization + ClassificationCategories: 8 (Sports, Politics, Finance, Medical, Technology, Culture, Religion, Accidents) AraMuST aims to support Arabic Natural Language Processing (NLP) research in areas such as: Abstractive text summarization Topic and news classification Multi-task learning and transfer learning Model evaluation and benchmarking All data sources were pre-existing open-access corpora, reprocessed and aligned for research reproducibility and academic use.

AraMuST是一款面向文本摘要与主题分类多任务学习的大规模统一阿拉伯语数据集。它整合并规范化了XL-Sum、SANAD及NADCG三大公开可用资源,基于统一标准化架构进行对齐整合。 该数据集提供高质量的文本-摘要配对数据,以及覆盖8大主流类别的统一主题标签,涵盖体育、政治、金融、科技、医疗、文化、宗教与事故领域。 为确保数据可靠性并避免数据泄露,本数据集经过了严格的质量筛选与多阶段验证流程。通过MD5哈希与基于嵌入的余弦相似度方法,检测并移除了不同数据源间的重复样本。 AraMuST的每条数据均遵循统一的JSONL格式规范,结构定义如下:{"task": "分类任务|摘要任务|分类+摘要联合任务", "text": 字符串类型, "summary": 字符串类型|空值(NA), "label": 字符串类型|空值(NA), "source": 字符串类型} 数据集构成:总样本量:2,143,833条;支持任务:摘要任务、分类任务及分类+摘要联合任务;类别覆盖:8大类,分别为体育、政治、金融、医疗、科技、文化、宗教与事故。 AraMuST旨在为阿拉伯语自然语言处理(Natural Language Processing,简称NLP)相关研究提供支撑,覆盖以下研究方向: - 抽象式文本摘要 - 主题与新闻分类 - 多任务学习与迁移学习 - 模型评估与基准测试 本数据集所用的全部数据源均为已公开的开放获取语料库,经重新处理与对齐后,可保障研究可复现性并满足学术使用需求。

提供机构:
Zenodo
创建时间:
2025-10-17
二维码
社区交流群
二维码
科研交流群
商业服务