遇见数据集

premio-ai/TheArabicPile_Medical

收藏
Hugging Face2024-03-21 更新2024-06-22 收录
官方服务:

资源简介:

--- language: - ar license: cc-by-nc-4.0 task_categories: - text-generation dataset_info: - config_name: dedup features: - name: text dtype: string splits: - name: train num_bytes: 19198059 num_examples: 32016 download_size: 9223336 dataset_size: 19198059 - config_name: default features: - name: text dtype: string splits: - name: train num_bytes: 34717649 num_examples: 53058 download_size: 11091835 dataset_size: 34717649 configs: - config_name: dedup data_files: - split: train path: dedup/train-* - config_name: default data_files: - split: train path: data/train-* --- # The Arabic Pile ![image/png](https://cdn-uploads.huggingface.co/production/uploads/64da0fd923557cdce3e514c3/J0oY67lVvecV75SOlWpjc.png) ## Introduction: The Arabic Pile is a comprehensive dataset meticulously designed to parallel the structure of The Pile and The Nordic Pile. Focused on the Arabic language, the dataset encompasses a vast array of linguistic nuances, incorporating both Modern Standard Arabic (MSA) and various Levantine, North African, and Egyptian dialects. Tailored for the training and fine-tuning of large language models, the dataset consists of 13 subsets, each uniquely crafted to cater to different linguistic domains. ## The Medical Subset: This dataset has a collection of all medical data collected on the interent for the Arabic language. The subset is quite limited and showcases the limitations in the Arabic content. ## Other Subsets: 1. premio-ai/TheArabicPile 2. premio-ai/TheArabicPile_Web 3. premio-ai/TheArabicPile_Lyrics 4. premio-ai/TheArabicPile_Reviews 5. premio-ai/TheArabicPile_Dialects 6. premio-ai/TheArabicPile_Mathematics 7. premio-ai/TheArabicPile_Conversational 8. premio-ai/TheArabicPile_Articles 9. premio-ai/TheArabicPile_Poetry 10. premio-ai/TheArabicPile_Medical 11. premio-ai/TheArabicPile_Miscellaneous 12. premio-ai/TheArabicPile_SocialMedia 13. premio-ai/TheArabicPile_Translations 14. premio-ai/TheArabicPile_Books These subsets serve distinct purposes, ranging from mathematical content to conversational dialogue, medical texts, and more. Notably, there's a dedicated subset, "premio-ai/TheArabicPile_SocialMedia," emphasizing the inclusion of language commonly found in social media contexts. ## Dataset Description * Curated by: Premio.AI team * Language(s) (NLP): Arabic, multiple languages on the translation dataset. * License: CC BY-NC 4.0 Deed - Non Commercial. * For any commercial uses or licensing, please contact mo@premio.ai. ## Data Structure The datasets are divided into two main subsets: 1. Original Subset: The raw data as collected from sources, without modifications. 2. Deduplication Subset: A filtered and cleaned version, enhancing usability for large language models by reducing redundancy and noise. The Arabic Pile extends an invitation not only for training and fine-tuning large language models but also for diverse applications across linguistic domains. Whether for research, analysis, or other linguistic endeavors, The Arabic Pile stands as a rich resource for the exploration of Arabic language intricacies. ## Data Collection Please refer to the paper for more details on our data collection procedures. ## Data Format The dataset has one single column called text. The text should contain the required meta data and the body combined. This was done to make sure that it will be a good fit for direct training or fine-tuning of large language models. Please note that the meta data might require to be repeated if your training context window won’t fit the entire body of text. ## Potential Bias As with any large-scale dataset, The Arabic Pile is not immune to potential biases that may influence the training and performance of language models. It's crucial to transparently address these biases to ensure responsible usage and interpretation of the dataset. Here are some considerations: 1. Dialectal Imbalance: The dataset incorporates various Arabic dialects, with a focus on Levantine, North African, and Egyptian variants. However, there might be variations in the representation of these dialects, potentially leading to an imbalance in the training data. 2. Source Influence: Bias may arise from the sources of the original data. The dataset collects information from diverse platforms and domains, and biases inherent in those sources could transfer to the dataset. 3. Social Media Context: Some of our datasets have language from social media platforms and online platforms. This subset may introduce biases inherent in online discourse, such as informal language, colloquial expressions, and potential subjectivity in politics, religion or culture. 4. Genre and Domain Bias: Different subsets cater to distinct linguistic domains, such as medical texts, poetry, reviews, and more. Each domain carries its own linguistic characteristics, potentially leading to biases based on the genres represented. ## License Information for The Arabic Pile: No Commercial Use The Arabic Pile is released under the Creative Commons Attribution-NonCommercial 4.0 International License (CC BY-NC 4.0). This license is designed to facilitate the open sharing and collaboration of the dataset while ensuring responsible and non-commercial usage. Key Points of the License: * Attribution (BY): Users are free to share, adapt, and build upon the dataset, even commercially, as long as they provide appropriate attribution to the dataset creators. * Non-Commercial (NC): The dataset may not be used for commercial purposes. Any use for commercial gain requires explicit permission from the dataset creators. * No Additional Restrictions: The license allows for maximum freedom of use, provided the terms of attribution and non-commercial use are adhered to. How to Cite: When using The Arabic Pile in your work, please include a proper citation to acknowledge the dataset creators. A recommended citation can be found in the model card for easy reference. License Deed: For a comprehensive understanding of the terms and conditions, please refer to the CC BY-NC 4.0 License Deed. By adopting this license, we aim to foster a collaborative and open environment for the exploration and advancement of Arabic language understanding and natural language processing. ## Citation When utilizing The Arabic Pile in your research, development, or other projects, we kindly request that you cite the dataset using the following format: @article{alrefaie2024arabicpile, author = {Mohamed Taher Alrefaie, Mahmoud Ibrahim Barbary, Ahmed Yasser Hassanein, Shiref Khaled Elhalawany, Karim Ashraf Elsayed, Ahmed Yasser }, title = {The Arabic Pile: A Large Scale Dataset of Diverse Text for Large Language Modeling}, year = {2024}, url = {https://huggingface.co/datasets/premio-ai/TheArabicPile} }

提供机构:
premio-ai
原始信息汇总

数据集概述

数据集信息

  • 语言: 阿拉伯语
  • 许可证: CC BY-NC 4.0(非商业用途)
  • 任务类别: 文本生成

配置详情

配置名称: dedup

  • 特征:
    • 名称: text
    • 数据类型: string
  • 分割:
    • 名称: train
    • 字节数: 19198059
    • 样本数: 32016
  • 下载大小: 9223336
  • 数据集大小: 19198059

配置名称: default

  • 特征:
    • 名称: text
    • 数据类型: string
  • 分割:
    • 名称: train
    • 字节数: 34717649
    • 样本数: 53058
  • 下载大小: 11091835
  • 数据集大小: 34717649

数据文件

  • 配置名称: dedup
    • 分割: train
    • 路径: dedup/train-*
  • 配置名称: default
    • 分割: train
    • 路径: data/train-*

数据集描述

  • 数据集名称: The Arabic Pile
  • 创建者: Premio.AI 团队
  • 语言: 阿拉伯语,翻译数据集包含多种语言
  • 许可证: CC BY-NC 4.0(非商业用途)
  • 商业用途: 请联系 mo@premio.ai

数据结构

  • 原始子集: 从来源收集的原始数据,未经修改。
  • 去重子集: 经过过滤和清洗的版本,通过减少冗余和噪声提高大型语言模型的可用性。

数据格式

  • 数据列: text
  • 文本内容: 包含所需的元数据和主体,以便直接用于大型语言模型的训练或微调。

潜在偏差

  • 方言不平衡: 数据集包含多种阿拉伯方言,但可能存在方言表示的差异。
  • 来源影响: 数据集从不同平台和领域收集信息,可能存在来源固有的偏差。
  • 社交媒体上下文: 部分数据集包含来自社交媒体平台的语言,可能引入在线讨论中的偏差。
  • 类型和领域偏差: 不同子集服务于不同的语言领域,每个领域都有其语言特征,可能导致基于类型的偏差。

许可证信息

  • 许可证: CC BY-NC 4.0(非商业用途)
  • 主要条款:
    • 署名 (BY): 用户可以自由分享、改编和构建数据集,只要提供适当的归属。
    • 非商业 (NC): 数据集不得用于商业目的。
    • 无额外限制: 只要遵守署名和非商业用途的条款,许可证允许最大程度的自由使用。

引用

  • 引用格式:

    @article{alrefaie2024arabicpile, author = {Mohamed Taher Alrefaie, Mahmoud Ibrahim Barbary, Ahmed Yasser Hassanein, Shiref Khaled Elhalawany, Karim Ashraf Elsayed, Ahmed Yasser }, title = {The Arabic Pile: A Large Scale Dataset of Diverse Text for Large Language Modeling}, year = {2024}, url = {https://huggingface.co/datasets/premio-ai/TheArabicPile} }

搜集汇总
数据集介绍
premio-ai/TheArabicPile_Medical 数据集图片
构建方式
在阿拉伯语自然语言处理领域,高质量医疗文本语料库的匮乏一直是制约领域模型发展的瓶颈。TheArabicPile_Medical数据集作为The Arabic Pile项目的一个子集,专门聚焦于阿拉伯语医疗文本的收集与整理。该数据集通过系统性地从互联网上采集公开的阿拉伯语医疗相关内容构建而成,涵盖现代标准阿拉伯语及多种方言变体。数据构建过程中,团队保留了原始采集的原始子集(default),同时提供了经过去重与清洗处理的去重子集(dedup),后者通过移除冗余与噪声数据,显著提升了语料质量,使其更适配大规模语言模型的训练需求。
使用方法
数据集以单一的text字段存储每条样本,将元数据与正文内容合并,这种设计使其可直接用于大语言模型的预训练或微调流程。用户可根据需求选择default或dedup两种配置,后者更适合追求高数据质量的场景。使用时需注意,若训练上下文窗口无法容纳完整文本,可能需要重复元数据部分以保证信息完整性。数据集采用CC BY-NC 4.0非商业许可协议,仅限学术研究用途,商业使用需联系版权方获取授权。
背景与挑战
背景概述
在自然语言处理领域,大规模、高质量的数据集是推动语言模型发展的基石,然而阿拉伯语因其复杂的方言体系与数字资源匮乏,长期面临语料库构建的瓶颈。2024年,由Premio.AI团队(核心成员包括Mohamed Taher Alrefaie、Mahmoud Ibrahim Barbary等)发布的TheArabicPile_Medical数据集,旨在填补阿拉伯语医学文本资源的空白。作为The Arabic Pile项目下的一个子集,该数据集专注于收集互联网上的阿拉伯语医学内容,涵盖现代标准阿拉伯语及黎凡特、北非、埃及等多种方言,为训练和微调大型语言模型提供专业领域的语料支撑。其创建不仅回应了阿拉伯语在医疗、科研等垂直领域的数据稀缺问题,更通过结构化设计(如去重版本)提升了数据可用性,对推动阿拉伯语自然语言处理在医学信息抽取、临床决策支持等方向的研究具有里程碑意义。
当前挑战
该数据集面临的首要挑战是领域内的数据稀缺性,其README文件明确指出医疗子集规模有限(去重后仅约3.2万条样本),这直接反映了阿拉伯语医学内容在互联网上的分布不足,限制了模型对复杂医学概念的泛化能力。其次,构建过程中需应对方言与标准语混杂带来的标注歧义,以及医疗术语在不同方言中的表达差异,增加了数据清洗和归一化的难度。此外,潜在偏见问题不容忽视:数据来源平台可能引入地域或机构偏好,导致疾病谱系或治疗方案的覆盖不均;同时,医学文本中隐含的文化与伦理敏感性要求模型在训练时需谨慎处理,以避免生成误导性建议。最后,非商业许可协议(CC BY-NC 4.0)虽保障了开放研究,却限制了其在商业医疗应用中的转化,构成了数据利用与伦理之间的张力。
常用场景
经典使用场景
在自然语言处理与医学信息学交叉领域,TheArabicPile_Medical数据集为阿拉伯语医学文本的预训练与微调提供了稀缺的高质量语料。该数据集汇聚了互联网上公开的阿拉伯语医学文献、临床记录及健康科普内容,涵盖现代标准阿拉伯语与多种方言表达,尤其适合用于构建面向阿拉伯语医疗场景的大语言模型。研究者可基于该数据集进行领域自适应的语言模型训练,提升模型在医学命名实体识别、临床文本摘要、症状-疾病关系抽取等下游任务中的表现。由于阿拉伯语医学资源长期匮乏,此数据集填补了语料空白,成为推动阿拉伯语医学自然语言处理研究的基础性资源。
解决学术问题
该数据集直面阿拉伯语医学自然语言处理中语料稀缺与方言多样性带来的学术挑战。传统医学NLP研究多集中于英语等资源丰富语言,而阿拉伯语因方言分歧、术语不统一及公开医学数据有限,导致模型在诊断辅助、药物信息抽取等任务上表现欠佳。TheArabicPile_Medical通过系统收集并去重处理医学文本,缓解了训练数据不足的瓶颈,使研究者得以探索跨方言医学语言理解、低资源场景下的迁移学习,以及医学本体与阿拉伯语词汇的语义对齐等问题。其发布推动了阿拉伯语医学信息学的实证研究,为构建公平、包容的医疗AI系统奠定了数据基础。
实际应用
在现实应用中,该数据集可赋能阿拉伯语地区的智能医疗系统。基于此数据集训练的模型能辅助医生进行电子病历的自动解析与结构化,例如从非结构化的临床笔记中提取诊断代码、药物名称及检验结果。同时,模型可支撑面向患者的阿拉伯语健康问答系统,为缺乏医疗资源的偏远地区提供初步的疾病咨询与用药指导。此外,在公共卫生监测领域,该数据集可用于训练分析社交媒体中与疾病相关的阿拉伯语讨论,实现疫情趋势的早期预警。这些应用不仅提升了医疗服务的效率,还降低了语言障碍对阿拉伯语使用者获取健康信息的限制。
数据集最近研究
最新研究方向
在阿拉伯语自然语言处理领域,大规模语料库的构建与多方言兼容性成为前沿焦点。TheArabicPile_Medical作为The Arabic Pile项目中的医疗子集,聚焦于阿拉伯语医学文本的收集与整理,其数据来源涵盖现代标准阿拉伯语及黎凡特、北非、埃及等地方言,直面当前阿拉伯语医疗语料稀缺的挑战。该子集虽规模有限,却精准指向低资源语言在专业领域(如医学)的瓶颈问题,为后续研究提供了针对性的基础资源。结合The Arabic Pile整体框架,该数据集不仅支持大语言模型在医疗问答、临床文本生成等任务上的微调,还通过去重版本优化了数据质量,助力缓解方言不平衡与领域偏见。这一方向呼应了全球对多语言AI公平性的关注,尤其在阿拉伯世界数字化转型加速的背景下,为构建包容性医疗NLP系统奠定了关键基石。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务