Latin-transliterated Ottoman Turkish Corpus (LATOC)
收藏资源简介:
LATOC Corpus LATOC (Latin-transliterated Ottoman Turkish Corpus) includes 143 Ottoman Turkish books, 13,252,350 words, written between the 15th and 20th centuries. The books were transliterated by domain experts and publicly shared on the Internet. The books in the corpus were automatically structured via a rule‑based approach and manually checked. Due to the copyright restrictions, this repository does not have any raw text; however, it guides you to download the files, convert them into structured XML files, and process them on your computer to get LATOC. The pipeline provided here can be extended to new sources by simply acquiring the PDF documents and, if necessary, updating the `CERTAIN_RULES` variable in 2_preprocessing.py and the files `controversy_cache.json` and `exclude_pages.txt`. Corpus Overview The corpus has more than 13 million words from 143 works written between the 15th and 20th centuries. While the pipeline standardizes these texts into the IJMES format, some inconsistencies, such as normalization of spelling, may persist. These are primarily inherited from the transliteration provided by the domain experts. This work does not apply any normalization to the spelling. Arabic/Persian characters in the documents are filtered out, and only the Latin-transliterated text is preserved. Each document is split into pages. Each page is divided into three segments: paragraph, comprising the main text, title, and footnote. The final XML files also provide the coordinates of the regions on the page, like "bbox="277.0,711.1,651.2,765.0". Data Files Work‑level metadata (`LATOC_metadata_sample.csv`) This is a sample of the metadata with further information. It provides metadata for 36 Dîvân works. Note that this metadata is from the previous version of LATOC, which is the reason why it has only 36 works. In the future, this scheme will be expanded to all works in LATOC. Each Dîvân work is accompanied by: - `file_name` - `work_name` (title of the Dîvân) - `pen_name` (mahlas) - `real_name` - `viaf` - `century` - `gender` - `rank` - e.g., “Sultan,” “Judiciary & Religious Office,” “High Bureaucracy/Military,” “Scholars & Sufi Orders,” “Civil Bureaucracy,” “Lay/Non‑official” Overall data statistics (`data_statistics.csv`) This file includes basic statistics such as the word count per document. It also provides links to some of the documents you can download for the work. Note that some works miss the URL links to download documents here. The document names in the column 'file' should be enough for users to find the document. Book‑level data Since the data was under copyright, this repository does not have it directly. However, you can download the data from _Yazma Eserler via either here or this webpage and then run the Python scripts as explained in this document to have the processed data on your device. Supplementary material (`controversy_cache.json` and `exclude_pages.txt`) `controversy_cache.json` includes the conversion rules for the problematic characters, which might be converted into more than one character in the IJMES chart. Since each document behaves differently, it provides the conversion rule based on the document. You can add a new rule here for your new documents. `exclude_pages.txt` has the page boundaries that should be deleted to remove the editorial preface, table of contents, and references, etc. You can enlarge this file if you add a new source to the data. Processing the data with Python scripts After downloading the files and storing them in a single folder, you should run the Python scripts 1_pdf_extractor.py, 2_preprocessing.py, and 3_xml_cleaner.py in turn. 2_preprocessing.py requires the supplementary file, `controversy_cache.json`. For 3_xml_cleaner, you need the supplementary document exclude_pages.txt. These files are prepared for 144 works presented in this dataset by the author. Usage Notes - The corpus can be utilized for **diachronic studies**; Yılandiloğlu (2025) demonstrated that poets adhered more accurately to the aruz meter over the centuries, reflected in rising conformity rates. - The sample metadata allows you to focus on specific ranks (e.g., “Sultan”) or gender. - Current work is focused on standardizing transliteration to the IJMES system and expanding the corpus further. Impact and Downstream Tasks This corpus was specifically curated and structured to support the development of Ottoman Turkish NLP resources. It has been used for: * **Large Language Models:** The structured data was used to train models, including: * Masked language model: [ota-roberta-base](https://huggingface.co/enesyila/ota-roberta-base) * State-of-the-art Named Entity Recognition model for Ottoman Turkish: [ota-roberta-base-ner](https://huggingface.co/enesyila/ota-roberta-base-ner) * A Universal Dependencies (UD) parser that tags with 91% accuracy and lemmatizes with 86% accuracy: [ota-ud-style](https://huggingface.co/enesyila/ota-ud-style) * **Annotated Treebank:** The dataset serves as the basis for [UD_Ottoman_Turkish-DUDU](https://github.com/UniversalDependencies/UD_Ottoman_Turkish-DUDU), currently the largest Ottoman Turkish corpus in the Universal Dependencies.
LATOC语料库(拉丁转写奥斯曼土耳其语语料库,Latin-transliterated Ottoman Turkish Corpus)包含143部奥斯曼土耳其语典籍,共计13252350个词元,创作时间跨度为15至20世纪。该语料库中的文本均由领域专家完成拉丁转写,并在互联网上公开共享。语料库中的典籍已通过基于规则的方法完成自动结构化,并经人工审核校验。 由于版权限制,本仓库未提供原始文本;但本仓库将指导用户自行下载相关文件,将其转换为结构化XML文件,并在本地计算机上进行处理以获取LATOC语料库。本文提供的处理流水线可扩展至新数据源:仅需获取PDF文档,必要时更新`2_preprocessing.py`中的`CERTAIN_RULES`变量,以及`controversy_cache.json`与`exclude_pages.txt`文件即可。 ## 语料库概览 本语料库收录了15至20世纪间的143部作品,总词量超1300万。尽管处理流水线将文本标准化为IJMES格式,但仍可能存在部分拼写规范化不一致的问题——此类问题主要源自领域专家提供的原始转写结果,本项目未对拼写进行任何规范化处理。文档中的阿拉伯语/波斯语字符已被过滤,仅保留拉丁转写文本。每份文档按页面拆分,每页分为三个段落区域:正文段落、标题与脚注。最终生成的XML文件还会提供页面内文本区域的坐标信息,例如`bbox="277.0,711.1,651.2,765.0"`。 ## 数据文件 ### 作品级元数据(`LATOC_metadata_sample.csv`) 本文件为元数据样本,包含36部《迪万》(Dîvân)作品的元数据信息。请注意,该元数据来自LATOC的旧版本,因此仅涵盖36部作品。未来该元数据方案将扩展至LATOC的全部作品。 每部《迪万》作品均对应以下字段: - `file_name`:文件名 - `work_name`(《迪万》作品标题) - `pen_name`(笔名,即mahlas) - `real_name`:真实姓名 - `viaf`:国际权威名称标识符(Virtual International Authority File) - `century`:创作世纪 - `gender`:创作者性别 - `rank`:社会阶层/身份等级,例如“苏丹”“司法与宗教职务”“高级官僚/军事阶层”“学者与苏菲教团”“文职官僚”“平民/非官方人士”。 ### 整体数据统计(`data_statistics.csv`) 本文件包含基础统计信息,例如单文档词量;同时提供了部分可下载作品的链接。请注意,部分作品未附带下载链接,不过`file`列中的文档名称已足够用户检索相关资源。 ### 单典籍级数据 由于版权限制,本仓库未直接提供该类数据。用户可通过此处或指定网页从_Yazma Eserler下载数据,随后按照本文档说明运行Python脚本,即可在本地设备上获取处理完成的数据。 ### 补充材料(`controversy_cache.json`与`exclude_pages.txt`) - `controversy_cache.json`:包含针对易混淆字符的转换规则——此类字符在IJMES字符表中可能对应多个转换结果。由于每份文档特性不同,本文件提供基于单文档的转换规则,用户可针对新文档添加自定义规则。 - `exclude_pages.txt`:收录了需删除的页面边界信息,用于移除编辑性前言、目录与参考文献等内容。若新增数据源,用户可扩展该文件内容。 ## Python脚本数据处理流程 下载文件并存储至单一文件夹后,请依次运行Python脚本`1_pdf_extractor.py`、`2_preprocessing.py`与`3_xml_cleaner.py`。其中`2_preprocessing.py`需依赖补充文件`controversy_cache.json`;`3_xml_cleaner.py`则需用到补充文档`exclude_pages.txt`。本项目已为本数据集收录的144部作品预置了上述文件。 ## 使用说明 - 本语料库可用于**历时研究**:Yılandiloğlu(2025)的研究表明,数个世纪以来诗人对阿鲁兹(aruz)格律的遵从度不断提升,这一趋势在文本的合规率增长中有所体现。 - 样本元数据可帮助用户聚焦于特定身份等级(如“苏丹”)或性别维度。 - 当前项目的工作重点为将转写结果标准化至IJMES体系,并进一步扩充语料库规模。 ## 应用价值与下游任务 本语料库专为支撑奥斯曼土耳其语自然语言处理(Natural Language Processing,NLP)资源开发而构建与结构化,已被应用于以下场景: - **大语言模型(Large Language Model)**:结构化数据已用于训练多款模型,包括: * 掩码语言模型:[ota-roberta-base](https://huggingface.co/enesyila/ota-roberta-base) * 奥斯曼土耳其语当前最优命名实体识别模型:[ota-roberta-base-ner](https://huggingface.co/enesyila/ota-roberta-base-ner) * 通用依存关系(Universal Dependencies,UD)句法分析器:该分析器的标注准确率达91%,词形还原准确率达86%,对应模型为[ota-ud-style](https://huggingface.co/enesyila/ota-ud-style) - **标注树库**:本数据集是[UD_Ottoman_Turkish-DUDU](https://github.com/UniversalDependencies/UD_Ottoman_Turkish-DUDU)的构建基础,该语料库是目前规模最大的通用依存关系体系下的奥斯曼土耳其语语料库。



