遇见数据集

Latin-transliterated Ottoman Turkish Corpus (LATOC)

收藏
Zenodo2025-08-22 更新2026-05-26 收录
官方服务:

资源简介:

This corpus includes 36 Ottoman Turkish poem books, Dîvân, written between 15th and 19th centuries. The books were transliterated by domain experts and publicly shared on the Internet. The books in the corpus were automatically structured via a rule-based approach and manually checked. See the interface for free: https://koras0ff.github.io/latoc/ Century Titles Poems Type Token 15 4 1,284 38,895 167,455 16 15 10,804 95,750 790,931 17 5 1,042 27,091 97,221 18 5 1,843 57,286 243,909 19 7 2,867 47,534 271,591 Total 36 17,840 188,441 1,571,107 Each Dîvân work and poem is accompanied by the following fields: Work-level (LATOC_metadata.csv) file_name work_name (title of the Dîvân) pen_name (mahlas) real_name viaf century gender rank (e.g., “Sultan,” “Judiciary & Religious Office,” "High Bureaucracy/Military," “Scholars & Sufi Orders,” “Civil Bureaucracy,” and “Lay/Non-official”) Poem-level (LATOC.json) poem_id title (if available) meter (in aruz notation) text (line-by-line Latin transliteration) The corpus can be utilised for diachronic studies as Yılandiloğlu (forthcoming) demonstrated that poets partially adhered more accurately to the aruz meter over the centuries, reflected in rising conformity rates. Additionally, metadata can be leveraged to focus on a specific rank such as sultan or gender. While this version has inconsistencies in terms of transliteration, current work is focused on standardizing the corpus according to IJMES transliteration system and increasing the size of the corpus.

本语料库包含36部创作于15至19世纪的奥斯曼土耳其语迪万诗集(Dîvân)。所有诗集均由领域专家完成拉丁转写,并于互联网上公开共享。语料库内的诗集均通过基于规则的方法完成自动结构化,并经人工核验。 可通过以下接口免费获取:https://koras0ff.github.io/latoc/ ### 语料统计详情 | 世纪 | 标题数 | 诗歌数 | Type | Token | |------|--------|--------|------|-------| | 15 | 4 | 1,284 | 38,895 | 167,455 | | 16 | 15 | 10,804 | 95,750 | 790,931 | | 17 | 5 | 1,042 | 27,091 | 97,221 | | 18 | 5 | 1,843 | 57,286 | 243,909 | | 19 | 7 | 2,867 | 47,534 | 271,591 | | 总计 | 36 | 17,840 | 188,441 | 1,571,107 | 每一部迪万诗集及单首诗歌均附带以下字段: #### 作品层级(LATOC_metadata.csv) - 文件名 - 作品名称(迪万诗集的标题) - 笔名(玛赫拉斯(mahlas)) - 真实姓名 - viaf - 世纪 - 性别 - 身份层级(例如:"苏丹""司法与宗教职务""高级官僚/军事人员""学者与苏菲教团成员""文职官僚""平民/非公职人员") #### 诗歌层级(LATOC.json) - 诗歌ID - 诗歌标题(若有) - 格律(采用阿鲁兹(aruz)记谱法) - 文本(逐行拉丁转写版) 本语料库可用于历时性研究:正如Yılandiloğlu(即将发表)的研究所示,数百年来诗人对阿鲁兹格律的遵从度逐步提升,这一趋势体现在合规率的持续走高中。此外,依托元数据可聚焦特定身份层级群体,例如苏丹群体或按性别细分研究。 当前版本的转写仍存在不一致之处,目前的后续工作正致力于按照《国际中东研究期刊》(IJMES)转写规范对语料库进行标准化,并扩充语料规模。

提供机构:
Zenodo
创建时间:
2025-07-20
二维码
社区交流群
二维码
科研交流群
商业服务