遇见数据集

BALT: Babylonian Administrative and Legal Texts on Oracc

收藏
Zenodo2026-03-17 更新2026-05-29 收录
官方服务:

资源简介:

This repository contains a copy of the data published as the project "BALT: Babylonian Administrative and Legal Texts" on the Open Richly Annotated Cuneiform Corpus (Oracc). The project contains 2,990 Babylonian administrative and legal texts from the Neo-Babylonian, Persian, and Hellenistic periods (c. 626–93 BCE). Stemming from the Eanna and Ebabbar temple archives in Uruk and Sippar and from private archives in Sippar, Babylon, Borsippa, Nippur, and Uruk, these texts give a picture of the administration and daily economic activities at ancient Babylonian cult centers and private households. These texts have been transliterated by a number of scholars specializing in first millennium cuneiform sources. The great majority of the texts were transliterated by the late János Everling, a Hungarian scholar who pioneered the practice of making cuneiform transliterations available online. His translations cover the texts published in AnOr 8, CT 49, GCCI 1–2, Nbk, TuM 2/3, UCP 9/1, UCP 9/3, UCP 9/12, VS 3, and YOS 17. Yuval Levavi and Caroline Waerzeggers provided transliterations and translations for the texts published in Administrative Epistolography in the Formative Phase of the Neo-Babylonian Empire (dubsar 3, 2018) and Marduk-rēmanni: Local Networks and Imperial Politics in Achaemenid Babylonia (OLA 233, 2014). Within the context of Oracc, the BALT project is special in that the lemmatizations, part of speech tags, normalizations, and sense tags were not done manually but rather semi-automatically, i.e. with the help of trainable language models. These models are largely but not completely accurate, hence certain words in the BALT corpus are not lemmatized or may have their lemma, normalization, POS tag, or sense wrong. According to our evaluation, about 93% of lemmas, 96% of POS tags, and 70% of normalizations are correct. Scripts.zip contains scripts used for converting CoNLL-U files to Oracc ATF. The BALT project was based at the Centre of Excellence in Ancient Near Eastern Empires, hosted at the University of Helsinki and funded by the Research Council of Finland (decision nos. 312051, 336673, and 352747). We thank Yuval Levavi and Caroline Waerzeggers for their permission to use their work in BALT. János Everling's legacy data is published to honor his pioneering work in making transliterated cuneiform texts available online. This semi-automatically lemmatized online edition of the texts has been created by Tero Alstola, Aleksi Sahala, Jonathan Valk, and Matthew Ong. Linda Leinonen, Matias Sakko, Senja Salmi, and Repekka Uotila assisted in cleaning the data and creating metadata. We wish to thank Kathleen Abraham, Michael Jursa, and Shai Gordin for giving us access to NaBuCCo metadata for certain texts. We also thank Niek Veldhuis and Heidi Jauhiainen for their help at various stages of the project. We are grateful to the Oracc steering committee and other developers of Oracc for providing us with the digital platform to publish this annotated text corpus. For further information on the dataset, see Alstola, T., Sahala, A., Valk, J., & Ong, M. (2026). Semi-Automatic Annotation of Babylonian Cuneiform Texts. Journal of Open Humanities Data, 12(41). https://doi.org/10.5334/johd.494.

本仓库收录了开放丰富注释楔形文字语料库(Open Richly Annotated Cuneiform Corpus, Oracc)上发布的“BALT:巴比伦行政与法律文献”项目的数据集副本。该项目包含2990篇来自新巴比伦、波斯以及希腊化时期(约公元前626年—公元前93年)的巴比伦行政与法律文献。这批文献源自乌鲁克的埃安娜(Eanna)与埃巴巴尔(Ebabbar)神庙档案、西帕尔的同类神庙档案,以及西帕尔、巴比伦、波尔西帕、尼普尔和乌鲁克的私人档案,生动展现了古代巴比伦宗教中心与私人家庭的行政运作模式与日常经济活动。 这些文献已由多位研究公元前一千纪楔形文字(cuneiform)资料的学者完成转写。其中绝大多数文献由已故匈牙利学者亚诺什·埃夫林(János Everling)完成转写,他率先开创了楔形文字转写文本线上公开传播的先河。他的译释内容涵盖《AnOr 8》《CT 49》《GCCI 1–2》《Nbk》《TuM 2/3》《UCP 9/1》《UCP 9/3》《UCP 9/12》《VS 3》与《YOS 17》中收录的文献。尤瓦尔·莱瓦维(Yuval Levavi)与卡罗琳·韦尔泽格斯(Caroline Waerzeggers)则为《新巴比伦帝国形成期的行政书信》(dubsar 3, 2018)与《马尔杜克·雷曼尼:阿契美尼德巴比伦的地方网络与帝国政治》(OLA 233, 2014)中收录的文献提供了转写与译释服务。 在Oracc的框架下,BALT项目的独特之处在于其词形还原(lemmatization)、词性标注(part-of-speech tagging)、规范化处理(normalization)以及语义标注(sense tagging)并非完全依靠人工完成,而是采用半自动化流程,即借助可训练语言模型(trainable language model)实现。此类模型的准确率较高但并非完美无缺,因此BALT语料库中的部分词汇未完成词形还原,或存在词元、规范化形式、词性标注或语义标注错误。根据我们的评估,约93%的词元、96%的词性标注以及70%的规范化形式均准确无误。 Scripts.zip压缩包中包含用于将CoNLL-U格式文件转换为Oracc ATF格式的脚本文件。 BALT项目依托赫尔辛基大学主办的古代近东帝国卓越研究中心开展,由芬兰研究理事会资助(资助编号:312051、336673与352747)。 我们谨此感谢尤瓦尔·莱瓦维与卡罗琳·韦尔泽格斯允许我们在BALT项目中使用其研究成果。亚诺什·埃夫林的遗留数据公开发布于此,以纪念他在推动楔形文字转写文本线上传播方面做出的开创性贡献。 本套采用半自动化词形还原流程的线上文献版本由特罗·阿尔斯托拉(Tero Alstola)、阿莱克西·萨哈拉(Aleksi Sahala)、乔纳森·瓦尔克(Jonathan Valk)与马修·翁(Matthew Ong)主创。琳达·莱伊宁(Linda Leinonen)、马蒂亚斯·萨科(Matias Sakko)、森娅·萨尔米(Senja Salmi)与雷佩卡·乌蒂拉(Repekka Uotila)协助完成了数据清理与元数据创建工作。 我们感谢凯瑟琳·亚伯拉罕(Kathleen Abraham)、迈克尔·尤尔萨(Michael Jursa)与沙伊·戈尔丁(Shai Gordin)允许我们获取部分文献的NaBuCCo元数据。我们同样感谢尼克·维尔德休伊斯(Niek Veldhuis)与海蒂·乔海伊宁(Heidi Jauhiainen)在项目各阶段提供的帮助。我们亦衷心感谢Oracc指导委员会与Oracc其他开发者为我们提供发布此注释文本语料库的数字平台。 如需了解该数据集的更多信息,请参阅:Alstola, T., Sahala, A., Valk, J., & Ong, M. (2026). Semi-Automatic Annotation of Babylonian Cuneiform Texts. *Journal of Open Humanities Data*, 12(41). https://doi.org/10.5334/johd.494.

提供机构:
Zenodo
创建时间:
2025-12-02
二维码
社区交流群
二维码
科研交流群
商业服务