遇见数据集

A hybrid approach to the small unannotated corpus-based language comparison and its application to the Old East Slavic charters - Supplementary material 2 (Modern East Slavic)

收藏
Zenodo2024-12-01 更新2026-05-26 收录
官方服务:

资源简介:

Modern East Slavic dialects (Belogornoje, Megra, Zialionka) General description A set of modern East Slavic Belogornoje, Megra and Zialionka small territorial lects subcorpora. Megra is an autochthonous (Barannikova, 2005) Northern Russian small territorial lect (Kryuchkova and Goldin, 2011). Belogornoje is a late settlement (Barannikova, 2005) Central Russian small territorial lect (Kryuchkova and Goldin, 2011). Zialionka is an autochthonous Northern Belarusian lect, radically different from Belogornoje and Megra by most of the key isoglosses within the Eastern part of the Slavic continuum. Sources Both Megra and Belogornoje texts originate from the Saratov dialectological corpus (Kryuchkova and Goldin, 2011). These are manually transcribed interviews with dialect speakers, mostly on the slice-of-life, rarely touching the topic of religion, recorded during the field trips of Saratov State University from 1980 to 2019. They possess some tagging, but for the purpose of clear cross-evaluation, the experiments do not use this information. The transcription is phonemic, faithful to the dialect features, and remains untouched in the experiments. Zialionka texts are also phonemically transcribed and untouched in experiments, they come from the Polack ethnographic collection (Lobač, 2011). The main genre is folklore tales, collected by transcribing interviews with small territorial lects speakers during the field trips of Polack State University (Belarus) from 1992 to 2010 years. There are no traces of notable phonetic irregularities within the texts. Unfortunately, there is no way to reliably establish it, as there are no available original recordings. The data statement is available among the downloadable files. How-to This section contains the tutorials that allow to use this data with the intended pipelines. Corpus-based distance measurement package The source code for package is available here, the manual is available in the README section of the repository. To use this dataset for the measurement of distance between Belogornoje, Megra and Zialionka lects, and their subsequent clusterisation, following steps should be completed: Download the Jupyter notebook that streamlines the package use. Download the dataset. Put the dataset into a selected folder on your computer (make sure there are no other files within this folder). Insert the path to the directory into CONTENT_DIR variable in the Jupyter notebook. Run the notebook, adjusting the parameters, if necessary.

### 现代东斯拉夫语支方言(别洛戈尔诺耶(Belogornoje)、梅格拉(Megra)、兹亚利翁卡(Zialionka)) #### 概况 本数据集为现代东斯拉夫语支下别洛戈尔诺耶、梅格拉与兹亚利翁卡三类小型地域语类(lect)的子语料库集合。其中,梅格拉(Megra)是一种本土(autochthonous,Barannikova, 2005)北方俄语小型地域语类(Kryuchkova与Goldin, 2011);别洛戈尔诺耶(Belogornoje)为晚近定居形成的(Barannikova, 2005)中部俄语小型地域语类(Kryuchkova与Goldin, 2011);兹亚利翁卡(Zialionka)是本土北方白俄罗斯语语类,在斯拉夫语连续体东部区域的绝大多数核心同言线(isogloss)层面,与别洛戈尔诺耶和梅格拉存在显著差异。 #### 语料来源 梅格拉与别洛戈尔诺耶的语料均源自萨拉托夫方言学语料库(Kryuchkova与Goldin, 2011)。这些语料为萨拉托夫国立大学1980年至2019年田野调查期间录制的方言使用者访谈录音的人工转写文本,内容多为日常生活话题,极少涉及宗教主题。部分语料带有标注,但为开展清晰的跨评估实验,本研究未使用该标注信息。转写采用音位转写方式,忠实还原方言特征,实验中未对其进行任何修改。 兹亚利翁卡的语料同样采用音位转写方式,实验中未作修改,其来源为波拉茨克民族志馆藏(Lobač, 2011)。该语料的主要体裁为民间故事,由白俄罗斯波拉茨克国立大学1992年至2010年田野调查期间,对小型地域语类使用者进行访谈并转写收集而来。文本中未发现显著的语音不规则现象,但遗憾的是,由于无原始录音可供参考,无法对此进行可靠验证。数据说明文件可在可下载文件中获取。 #### 使用指南 本节包含将本数据集与预设流程配合使用的教程。 ##### 基于语料库的距离测算工具包 该工具包的源代码可在此处获取,使用手册详见代码仓库的README章节。 若需使用本数据集测算别洛戈尔诺耶、梅格拉与兹亚利翁卡语类间的距离,并对其进行后续聚类分析,请完成以下步骤: 1. 下载用于简化工具包使用流程的Jupyter Notebook文件; 2. 下载本数据集; 3. 将数据集放置于计算机中选定的文件夹内(请确保该文件夹内无其他文件); 4. 在Jupyter Notebook中将该目录路径赋值给CONTENT_DIR变量; 5. 运行该Notebook,必要时可调整参数。

提供机构:
Zenodo
创建时间:
2024-12-01
二维码
社区交流群
二维码
科研交流群
商业服务