SIMBA9657/haddas-tigrinya-corpus
收藏官方服务:
资源简介:
该数据集是一个单语提格里尼亚语报纸文本语料库,内容分割成文章主体。它适用于继续预训练或提格里尼亚语大语言模型的因果语言建模任务。数据来源于Haddas Eritrea报纸档案,经过提取、清洗、分割、翻译和标注等处理流程,包含2653行数据,模式包括id、text、char_count、topic、issue_date、source_pdf、page_start和page_end字段。数据集针对低资源语言场景设计,支持文本生成和掩码填充等NLP任务。
Monolingual Tigrinya newspaper text segmented into article bodies. Suitable for continued pretraining or causal language modeling of Tigrinya LLMs. Derived from the Haddas Eritrea newspaper archive: 63 PDF issues processed by the haddas-eritrea pipeline (extract -> clean -> segment -> translate -> label). Row count: 2653, with schema: id, text, char_count, topic, issue_date, source_pdf, page_start, page_end.
提供机构:
SIMBA9657


