sol-r/keilschrift-corpus
收藏资源简介:
# Keilschrift: Ancient Near Eastern Cuneiform Corpus A unified multilingual corpus of ancient Near Eastern cuneiform texts for machine learning, covering Sumerian, Akkadian, Old Persian, Elamite, Hittite, Hurrian, Urartian, and Aramaic — spanning 3400 BCE to 63 BCE. "Keilschrift" is German for "cuneiform" (literally "wedge-writing"). ## Statistics | | count | |---|---| | **corpus passages** | 214,049 | | **parallel pairs** | 170,199 | | **languages** | 8 | | **time span** | 3400 BCE – 63 BCE | ### Corpus by language | ISO 639-3 | language | family | count | period | |---|---|---|---|---| | `sux` | Sumerian | isolate | 182,621 | 3400 BCE – 100 CE | | `akk` | Akkadian | Semitic (East) | 21,929 | 2500 BCE – 100 CE | | `peo` | Old Persian | IE (Iranian) | 8,045 | 525 – 330 BCE | | `elx` | Elamite | isolate | 1,436 | 2600 – 360 BCE | | `xhu` | Hurrian | Hurro-Urartian | 12 | 2300 – 1200 BCE | | `hit` | Hittite | IE (Anatolian) | 2 | 1600 – 1178 BCE | | `xur` | Urartian | Hurro-Urartian | 2 | 860 – 590 BCE | | `arc` | Aramaic | Semitic (NW) | 2 | 1100 BCE – present | ### Corpus by source | source | passages | |---|---| | CDLI | 131,316 | | SumTablets | 82,339 | | ETCSL | 394 | ### Pairs by type | direction | type | count | |---|---|---| | sux → sux | cuneiform unicode → transliteration | 82,339 | | sux → sux | sign names → transliteration | 82,339 | | sux → eng | transliteration → translation | 4,455 | | akk → eng | transliteration → translation | 961 | | elx → eng | transliteration → translation | 72 | | peo → eng | transliteration → translation | 29 | | hit → eng | transliteration → translation | 2 | | xur → eng | transliteration → translation | 2 | ## Sources and citations ### SumTablets 82,339 Sumerian cuneiform tablets with three-way alignment: Unicode cuneiform glyphs, sign name sequences, and scholarly transliterations. Derived from ORACC transliterations mapped to Unicode representations. - **Citation:** Simmons, C. (2024). "SumTablets: A Transliteration Dataset of Sumerian Tablets." *Proceedings of the 1st Workshop on Machine Learning for Ancient Languages (ML4AL 2024)*, ACL Anthology. doi:[10.18653/v1/2024.ml4al-1.20](https://aclanthology.org/2024.ml4al-1.20/) - **License:** CC BY 4.0 - **URL:** https://huggingface.co/datasets/colesimmons/SumTablets ### ETCSL (Electronic Text Corpus of Sumerian Literature) 394 Sumerian literary compositions with transliterations and 380 English prose translations. Includes hymns, myths, epics (Gilgamesh, Enmerkar), laments, proverbs, debates, and royal hymns. - **Citation:** Black, J.A., Cunningham, G., Fluckiger-Hawker, E., Robson, E., and Zólyomi, G. (1998–2006). *The Electronic Text Corpus of Sumerian Literature.* Oxford: Faculty of Oriental Studies, University of Oxford. - **License:** CC BY-NC-SA 3.0 - **URL:** https://etcsl.orinst.ox.ac.uk/ - **Archive:** Oxford Text Archive, handle [20.500.12024/2518](https://ota.bodleian.ox.ac.uk/repository/xmlui/handle/20.500.12024/2518) ### CDLI (Cuneiform Digital Library Initiative) 131,316 cuneiform texts with ATF transliterations drawn from a catalogue of 353,000+ artifacts. Covers Sumerian, Akkadian, Old Persian, Elamite, Hittite, Hurrian, Urartian, and Aramaic. 12,559 texts include English translations. - **Citation:** Englund, R.K. et al. *Cuneiform Digital Library Initiative.* University of California, Los Angeles / Max Planck Institute for the History of Science, Berlin. - **License:** CC BY 4.0 - **URL:** https://cdli.earth/ - **Data:** https://github.com/cdli-gh/data ### ORACC (Open Richly Annotated Cuneiform Corpus) Not yet included in this release. ORACC provides curated, lemmatized cuneiform texts with scholarly translations from 140+ projects covering Sumerian, Akkadian, Hittite, and other languages. - **Citation:** Tinney, S. et al. *Open Richly Annotated Cuneiform Corpus.* University of Pennsylvania Museum of Archaeology and Anthropology. - **License:** CC BY-SA 3.0 - **URL:** https://oracc.museum.upenn.edu/ ## Schema **Corpus (monolingual passages):** | field | type | description | |---|---|---| | `id` | string | unique identifier (`{source}_{original_id}`) | | `text` | string | text content (transliteration) | | `text_type` | string | `transliteration` | | `language` | string | ISO 639-3 code | | `source` | string | `etcsl`, `cdli`, `sumtablets` | | `period` | string | historical period (e.g. "Ur III (ca. 2100-2000 BC)") | | `genre` | string | text genre (e.g. "literary", "administrative") | | `cdli_id` | string | CDLI P-number if available | | `work` | string | composition name for literary texts | | `word_count` | int | word/token count | **Pairs (parallel texts):** | field | type | description | |---|---|---| | `id` | string | unique identifier | | `source` | string | data source | | `language_a` | string | source language ISO 639-3 | | `language_b` | string | target language ISO 639-3 | | `text_a` | string | source text | | `text_b` | string | target text | | `text_type_a` | string | `cuneiform_unicode`, `sign_names`, or `transliteration` | | `text_type_b` | string | `transliteration` or `translation` | | `period` | string | historical period | | `genre` | string | text genre | | `cdli_id` | string | CDLI P-number | | `work` | string | composition name | ## Related projects - [CuneiML](https://github.com/taineleau/CuneiML) — cuneiform dataset with tablet photographs and Unicode transcriptions (Liang et al., 2023). doi:[10.5334/johd.151](https://openhumanitiesdata.metajnl.com/articles/10.5334/johd.151) - [CompVis cuneiform-sign-detection](https://github.com/CompVis/cuneiform-sign-detection-dataset) — image-based cuneiform sign detection dataset ## Acknowledgments This dataset builds on decades of work by assyriologists, sumerologists, and digital humanities scholars who created CDLI, ORACC, and ETCSL. Special thanks to Cole Simmons and the SumTablets project for making cuneiform Unicode mappings available in ML-ready format.
# Keilschrift:近东古代楔形文字语料库 本数据集为面向机器学习的近东古代楔形文字多语言统一语料库,涵盖苏美尔语、阿卡德语、古波斯语、埃兰语、赫梯语、胡里安语、乌拉尔图语与阿拉姆语,时间跨度为公元前3400年至公元前63年。 *注:Keilschrift为德语,意为“楔形文字”(字面含义为“楔形书写”)。* ## 统计数据 | | 数量 | |---|---| | **语料库篇章** | 214049 | | **平行语料对** | 170199 | | **覆盖语言** | 8种 | | **时间跨度** | 公元前3400年 – 公元前63年 | ### 按语言划分的语料库 | ISO 639-3代码 | 语言 | 语系 | 篇章数量 | 使用时期 | |---|---|---|---|---| | `sux` | 苏美尔语 | 孤立语 | 182621 | 公元前3400年 – 公元100年 | | `akk` | 阿卡德语 | 东闪米特语系 | 21929 | 公元前2500年 – 公元100年 | | `peo` | 古波斯语 | 印欧语系(伊朗语族) | 8045 | 公元前525年 – 公元前330年 | | `elx` | 埃兰语 | 孤立语 | 1436 | 公元前2600年 – 公元前360年 | | `xhu` | 胡里安语 | 胡里安-乌拉尔图语系 | 12 | 公元前2300年 – 公元前1200年 | | `hit` | 赫梯语 | 印欧语系(安纳托利亚语族) | 2 | 公元前1600年 – 公元前1178年 | | `xur` | 乌拉尔图语 | 胡里安-乌拉尔图语系 | 2 | 公元前860年 – 公元前590年 | | `arc` | 阿拉姆语 | 西北闪米特语系 | 2 | 公元前1100年 – 至今 | ### 按来源划分的语料库 | 数据来源 | 篇章数量 | |---|---| | CDLI(楔形文字数字图书馆倡议,Cuneiform Digital Library Initiative) | 131316 | | SumTablets | 82339 | | ETCSL(苏美尔文学电子文本语料库,Electronic Text Corpus of Sumerian Literature) | 394 | ### 按类型划分的平行语料对 | 语向 | 类型 | 数量 | |---|---|---| | sux → sux | 楔形文字Unicode字符 → 转写 | 82339 | | sux → sux | 符号名称 → 转写 | 82339 | | sux → eng | 转写 → 译文 | 4455 | | akk → eng | 转写 → 译文 | 961 | | elx → eng | 转写 → 译文 | 72 | | peo → eng | 转写 → 译文 | 29 | | hit → eng | 转写 → 译文 | 2 | | xur → eng | 转写 → 译文 | 2 | ## 数据来源与引用 ### SumTablets 本数据集包含82339篇苏美尔语楔形文字泥板文本,实现三重对齐:楔形文字Unicode字符、符号名称序列与学术转写结果。数据源自映射至Unicode表示形式的ORACC(开放富注释楔形文字语料库,Open Richly Annotated Cuneiform Corpus)转写内容。 - **引用格式**:Simmons, C. (2024). "SumTablets: A Transliteration Dataset of Sumerian Tablets." *Proceedings of the 1st Workshop on Machine Learning for Ancient Languages (ML4AL 2024)*, ACL Anthology. doi:[10.18653/v1/2024.ml4al-1.20](https://aclanthology.org/2024.ml4al-1.20/) - **许可协议**:CC BY 4.0(知识共享署名4.0国际许可协议) - **访问链接**:https://huggingface.co/datasets/colesimmons/SumTablets ### ETCSL(苏美尔文学电子文本语料库,Electronic Text Corpus of Sumerian Literature) 本数据集包含394篇苏美尔文学作品,涵盖转写内容与380篇英语散文译文,涉及赞美诗、神话、史诗(吉尔伽美什、恩美卡尔)、挽歌、谚语、辩论文本与王室赞美诗等体裁。 - **引用格式**:Black, J.A., Cunningham, G., Fluckiger-Hawker, E., Robson, E., and Zólyomi, G. (1998–2006). *The Electronic Text Corpus of Sumerian Literature.* Oxford: Faculty of Oriental Studies, University of Oxford. - **许可协议**:CC BY-NC-SA 3.0(知识共享署名-非商业性使用-相同方式共享3.0国际许可协议) - **访问链接**:https://etcsl.orinst.ox.ac.uk/ - **存档地址**:牛津文本档案馆,标识符[20.500.12024/2518](https://ota.bodleian.ox.ac.uk/repository/xmlui/handle/20.500.12024/2518) ### CDLI(楔形文字数字图书馆倡议,Cuneiform Digital Library Initiative) 本数据集包含131316篇楔形文字文本,采用ATF转写格式,源自超过353000件文物的目录,涵盖苏美尔语、阿卡德语、古波斯语、埃兰语、赫梯语、胡里安语、乌拉尔图语与阿拉姆语八种语言,其中12559篇文本附带英语译文。 - **引用格式**:Englund, R.K. et al. *Cuneiform Digital Library Initiative.* University of California, Los Angeles / Max Planck Institute for the History of Science, Berlin. - **许可协议**:CC BY 4.0(知识共享署名4.0国际许可协议) - **访问链接**:https://cdli.earth/ - **数据仓库**:https://github.com/cdli-gh/data ### ORACC(开放富注释楔形文字语料库,Open Richly Annotated Cuneiform Corpus) 本版本暂未收录ORACC数据。ORACC提供经过整理、词形还原的楔形文字文本,附带来自140余个项目的学术译文,涵盖苏美尔语、阿卡德语、赫梯语等多种语言。 - **引用格式**:Tinney, S. et al. *Open Richly Annotated Cuneiform Corpus.* University of Pennsylvania Museum of Archaeology and Anthropology. - **许可协议**:CC BY-SA 3.0(知识共享署名-相同方式共享3.0国际许可协议) - **访问链接**:https://oracc.museum.upenn.edu/ ## 数据结构 ### 单语文本语料库(单语篇章) | 字段 | 数据类型 | 描述 | |---|---|---| | `id` | 字符串 | 唯一标识符,格式为`{source}_{original_id}` | | `text` | 字符串 | 文本内容(转写形式) | | `text_type` | 字符串 | 文本类型,固定为`transliteration`(转写) | | `language` | 字符串 | ISO 639-3语言代码 | | `source` | 字符串 | 数据来源,可选值为`etcsl`、`cdli`、`sumtablets` | | `period` | 字符串 | 历史时期(例如“乌尔第三王朝(约公元前2100-2000年)”) | | `genre` | 字符串 | 文本体裁(例如“文学类”“行政类”) | | `cdli_id` | 字符串 | 若可用则为CDLI P编号 | | `work` | 字符串 | 文学作品的作品名称 | | `word_count` | 整数 | 单词/Token计数 | ### 平行文本语料对 | 字段 | 数据类型 | 描述 | |---|---|---| | `id` | 字符串 | 唯一标识符 | | `source` | 字符串 | 数据来源 | | `language_a` | 字符串 | 源语言ISO 639-3代码 | | `language_b` | 字符串 | 目标语言ISO 639-3代码 | | `text_a` | 字符串 | 源文本 | | `text_b` | 字符串 | 目标文本 | | `text_type_a` | 字符串 | 源文本类型,可选值为`cuneiform_unicode`(楔形文字Unicode字符)、`sign_names`(符号名称)或`transliteration`(转写) | | `text_type_b` | 字符串 | 目标文本类型,可选值为`transliteration`(转写)或`translation`(译文) | | `period` | 字符串 | 历史时期 | | `genre` | 字符串 | 文本体裁 | | `cdli_id` | 字符串 | CDLI P编号 | | `work` | 字符串 | 作品名称 | ## 相关项目 - [CuneiML](https://github.com/taineleau/CuneiML) — 包含楔形文字泥板照片与Unicode转写的楔形文字数据集(Liang等,2023),DOI:[10.5334/johd.151](https://openhumanitiesdata.metajnl.com/articles/10.5334/johd.151) - [CompVis楔形文字符号检测数据集](https://github.com/CompVis/cuneiform-sign-detection-dataset) — 基于图像的楔形文字符号检测数据集 ## 致谢 本数据集基于数十年来亚述学家、苏美尔学家与数字人文领域学者构建CDLI、ORACC与ETCSL的长期工作成果。特别感谢Cole Simmons与SumTablets项目团队,将楔形文字Unicode映射整理为适配机器学习的格式。



