somali-web-corpus
收藏资源简介:
SOMALI-WEB-CORPUS V1 是一个索马里语文本数据集,专门为支持索马里语的自然语言处理研究而构建。该数据集从多个在线来源(包括 BBC Somali、Shabeele media 和 SONNA(索马里国家通讯社)等网站)收集文本,并经过严格的清理和过滤流程。清理工作包括去除网页URL、电子邮件地址和无效顶级域名等无关信息,并对文本进行规范化处理。同时,通过去重操作丢弃了重复的段落,以确保数据内容的唯一性。数据集以 JSON 行文件(.jsonl)格式提供,其中每条记录仅包含一个名为 text 的字段,该字段存储一段经过清理的索马里语段落文本。该数据集的主要目的是用于训练和微调索马里语的语言模型(LLMs),适用于文本生成等自然语言处理任务。
SOMALI-WEB-CORPUS V1 is a Somali text dataset specifically constructed to support natural language processing research for the Somali language. The dataset collects text from multiple online sources, including websites such as BBC Somali, Shabeele media, and SONNA (Somali National News Agency), and undergoes a rigorous cleaning and filtering process. Cleaning involves removing irrelevant information like web URLs, email addresses, and invalid top-level domains, as well as normalizing the text. Additionally, duplicate paragraphs are discarded through deduplication to ensure the uniqueness of the data content. The dataset is provided in JSON Lines (.jsonl) format, where each record contains only one field named text, which stores a cleaned Somali paragraph. The primary purpose of this dataset is for training and fine-tuning Somali language models (LLMs) and is suitable for natural language processing tasks such as text generation.
数据集概述
数据集名称:Somali Web Corpus V1
语言:索马里语(so)
许可证:MIT
任务类别:文本生成
标签:索马里语、语料库、文本生成、大语言模型
数据集详情
- 格式:JSON lines 格式(.jsonl)
- 数据结构:每条记录包含一个
text字段,存储清洗后的段落文本 - 数据来源:文本提取自多个索马里语网站,包括 BBC Somali、Shabeele media、SONNA(索马里国家新闻局)等
- 数据清洗与过滤:
- 去重:移除重复的段落,确保数据唯一性
- 标准化:删除网页 URL、电子邮件地址以及无效的顶级域名
用途
该数据集专为训练和微调索马里语言模型(LLM)而设计,旨在支持索马里语的自然语言处理(NLP)研究。




