Northeastern Neo-Aramaic Corpus Data
收藏资源简介:
东北新阿拉姆语料库数据包含了一组非常多样的新阿拉姆语方言,这些方言直到现代仍在伊拉克北部、伊朗西北部和土耳其东南部的基督教和犹太社区中使用。这些是阿拉姆语的最后遗存之一,阿拉姆语是古代该地区的主要语言之一。该文本语料库由[Geoffrey Khan教授](https://www.ames.cam.ac.uk/people/professor-geoffrey-khan)及其团队收集的转录和录音文本组成,旨在保存这些日益濒危的语言。
The Northeastern Neo-Aramaic Corpus comprises a highly diverse collection of Neo-Aramaic dialects, which are still spoken today by Christian and Jewish communities in northern Iraq, northwestern Iran, and southeastern Turkey. These dialects represent some of the last vestiges of Aramaic, one of the principal languages of the ancient region. The text corpus consists of transcribed and recorded texts gathered by Professor Geoffrey Khan and his team, with the aim of preserving these increasingly endangered languages.
Northeastern Neo-Aramaic Corpus Data
数据集概述
Northeastern Neo-Aramaic Corpus Data 包含了一系列多样化的阿拉米语方言,这些方言直至现代仍在伊拉克北部、伊朗西北部和土耳其东南部的基督教和犹太社区中使用。这些方言是阿拉米语这一古代地区主要语言的最后存活的遗迹之一。
数据集内容
- nena_format - 描述NENA标记格式的文档。
- standards - NENA字母表和语言代码的标准,包括正则表达式模式。
- texts - 包含按版本和方言分类的NENA文本,采用NENA标记格式。
- parsed_texts - 包含所有解析为JSON层次结构的NENA文本。
- text_parser - 用于从NENA标记生成NENA JSON解析的SLY解析器。
- sources - 用于生成NENA文本的原始源材料。
数据集用途
该数据集旨在收集源文本,用于构建Text-Fabric中的完整文本语料库。该语料库将用于语言特征的分析和注释。




