SiDiaC-v.2.0
收藏资源简介:
SiDiaC-v.2.0是当前最大的僧伽罗语历时语料库,由斯里兰卡莫拉图瓦大学和信息技术研究院联合构建,覆盖公元5世纪至20世纪的文献。该语料库包含18.5万部文学作品共计24.1万词项,数据源自斯里兰卡国家图书馆的扫描文献,经谷歌Document AI OCR数字化后,通过多阶段处理流程解决格式错误、混合编码等问题。语料库采用双层分类体系,按虚构/非虚构进行主分类,并细分为宗教、历史等次级类别,为低资源语言僧伽罗语的历时语言演变研究及NLP任务提供重要资源。
SiDiaC-v.2.0 is the largest existing diachronic Sinhala corpus to date. It was jointly constructed by the University of Moratuwa and the Institute of Information Technology, Sri Lanka, covering documents spanning from the 5th century CE to the 20th century. This corpus includes 185,000 literary works with a total of 241,000 lexical items. The source data is derived from scanned documents held by the National Library of Sri Lanka, which were first digitized using Google Document AI OCR and then processed via a multi-stage workflow to address formatting errors, mixed encoding and other issues. The corpus adopts a two-tier classification system: it is primarily categorized into fictional and non-fictional works, with further subdivisions into secondary categories including religion, history and more. It serves as a critical resource for diachronic language evolution research and natural language processing (NLP) tasks targeting the low-resource language Sinhala.
SiDiaC-v.2.0: Sinhala Diachronic Corpus Version 2.0 数据集概述
数据集简介
SiDiaC-v.2.0 是一个经过整理的僧伽罗语历时语料库,包含僧伽罗语文学文本及相关资源。该数据集为每部作品提供了光学字符识别(OCR)输出和基本元数据。
核心内容
- 数据构成:包含僧伽罗语文学文本的原始PDF文件、最终OCR文本以及每部作品的机器可读元数据文件。
- 资源类型:文学文本、OCR输出、元数据。
数据结构
Books_PDF/目录:存放每部作品的原始源PDF文件。OCR_Final/目录:存放每部作品的最终OCR结果,每个子目录包含:metadata.json:包含基本的书目和处理元数据。<title>.txt:通过OCR提取的纯文本内容。
元数据示例
元数据文件(metadata.json)包含以下字段:
title:作品标题(僧伽罗语)。title_en:作品英文标题。author:作者(僧伽罗语)。author_en:作者英文名。genre:体裁/分类。issued_date:发行/出版日期。written_date:写作日期范围。ocr_confidence:OCR过程产生的启发式置信度分数。
重要说明
- 文件名和目录名可能包含僧伽罗语字符。
ocr_confidence是OCR过程的启发式评分,可能因作品而异。



