Leibniz-ZAS corpus of MAIN
收藏资源简介:
The presented dataset is part of the narrative corpus collected at Leibniz-Centre General Linguistics (Leibniz-ZAS). It contains transcriptions of oral narratives elicited with the <em>Multilingual Assessment Instrument for Narratives</em> (MAIN; read more here), developed as part of the LITMUS battery of tests in the framework of COST Action IS0804 <em>Language Impairment in a Multilingual Society: Linguistic Patterns and the Road to Assessment</em>. Narratives were elicited in the Russian, Turkish and German languages in the telling elicitation mode using two MAIN picture stories, Baby Birds and Baby Goats. The data were collected during two large-scale longitudinal studies conducted at ZAS in the framework of the <em>Berlin Interdisciplinary Network for Multilingualism (BIVEM)</em> and <em>Interdisciplinary Research Alliance (IFV)</em> projects (more information about the studies). The participants of the studies were Russian-German and Turkish-German bilingual children from different areas of Berlin. Their language development was closely documented every year from early kindergarten up to the end of the third grade of primary school (age 2;9 to 10;4 years). It is the longest and largest study of language development in bilingual children in Germany allowing for cross-sectional and longitudinal analyses from a cross-linguistic perspective. The narratives were audio recorded and transcribed in the standardized CHAT format (MacWhinney, 2000) using the CLAN program according to the CHILDES transcription rules for later analysis. The transcriptions can be used to analyze the narrative abilities of bilingual children on macro- and microstructural levels (more information can be found here). In total, the dataset contains 210 transcriptions of narratives from 29 participants (10 Russian-German bilingual children and 19 Turkish-German bilingual children), who were tested 5 times after the initial testing (pretest). The 5 testing points are therefore referred to as posttests: post1, post2, post3, post4, post5, post6 (this dataset does not contain data from post5, as oral narratives were not elicited at the end of the second grade). The corresponding age ranges at all testing points are given below for each part of the dataset. The dataset is divided into two parts, Russian-German and Turkish-German narrative corpus respectively. The narrative corpus of Russian-German bilingual children includes two folders with narratives elicited in Russian and German, at 5 testing points. Total number of transcriptions=100 Number of children=10 Total age range=2;9-10;4 Age range of children for narratives in Russian at each testing point: post 1: 2;9-4;3 (kindergarten) post 2: 3;9-5;2 (kindergarten) post 3: 4;9-6;1 (kindergarten) post 4: 6;9-7;6 (end of first grade) post 6: 8;7-9;10 (end of third grade) Age range of children for narratives in German at each testing point: post 1: 2;10-4;3 (kindergarten) post 2: 3;9-5;3 (kindergarten) post 3: 4;9-6;2 (kindergarten) post 4: 6;9-7;6 (end of first grade) post 6: 8;8-10;4 (end of third grade) The narrative corpus of Turkish-German bilingual children includes two folders. One folder contains narratives elicited in German at the earlier 3 testing points, which allows the analysis of early narrative development in one language. Total number of transcriptions=30 Number of children=10 Total age range=3;5-6;4 Age range of children for narratives in German at each testing point: post 1: 3;5-4;3 (kindergarten) post 2: 4;4-5;4 (kindergarten) post 3: 5;3-6;4 (kindergarten) Another folder contains narratives elicited in both languages, Turkish and German, at 4 testing points starting from post2 and allowing for the analysis of narrative development up to the third grade in both languages. Total number of transcriptions=80 Number of children=10 Total age range=3;10-9;9 Age range of children for narratives in Turkish at each testing point: post 2: 3;10-5;1 (kindergarten) post 3: 4;9-6;1 (kindergarten) post 4: 6;5-7;8 (end of first grade) post 6: 8;6-9;9 (end of third grade) Age range of children for narratives in German at each testing point: post 2: 4;1-5;4 (kindergarten) post 3: 5;1-6;4 (kindergarten) post 4: 6;6-7;8 (end of first grade) post 6: 8;5-9;8 (end of third grade) The files are named according to the following pattern: child’s code (letters refer to child’s first languages: r-Russian, t-Turkish), test (MAIN), story (bb=Baby Birds, bg=Baby Goats), language of elicitation (de/ru/tr), testing point (1=post1, 2=post2 etc.), and child’s age (year/month). Here is an example: r009_MAIN_bb_de_4_610.
本数据集隶属于莱布尼茨普通语言学研究中心(Leibniz-Centre General Linguistics, Leibniz-ZAS)采集的叙事语料库,包含采用**叙事多语言评估工具(Multilingual Assessment Instrument for Narratives, MAIN)**采集的口头叙事转写文本。该工具作为COST行动IS0804框架下LITMUS测试组合的一部分开发完成,COST行动IS0804的全称是“多语言社会中的语言障碍:语言模式与评估路径”(Language Impairment in a Multilingual Society: Linguistic Patterns and the Road to Assessment)。 叙事采集采用讲述诱导范式,使用两套MAIN图片故事:《幼鸟》(Baby Birds)与《幼羊》(Baby Goats),覆盖俄语、土耳其语与德语三种语言。 本数据集的采集工作依托莱布尼茨ZAS开展的两项大型纵向研究完成,相关研究隶属于**柏林多语言跨学科网络(Berlin Interdisciplinary Network for Multilingualism, BIVEM)**与**跨学科研究联盟(Interdisciplinary Research Alliance, IFV)**项目,研究详情可参见此处。 研究参与者为柏林不同区域的俄-德双语儿童与土-德双语儿童,研究人员每年对其语言发展进行详细记录,追踪时段从幼儿园早期直至小学三年级结束,年龄范围为2;9至10;4岁。该项研究是德国境内规模最大、追踪周期最长的双语儿童语言发展研究,支持跨语言视角下的横断与纵向分析。 所有叙事均进行音频录制,并采用**CHAT转写格式(CHAT)**,借助**CLAN程序(CLAN)**遵循**CHILDES转写规则(CHILDES)**完成转写,以供后续分析使用。此类转写文本可用于从宏观结构与微观结构层面分析双语儿童的叙事能力,相关详情可参见此处。 本数据集总计包含29名参与者的210份叙事转写文本,其中俄-德双语儿童10名,土-德双语儿童19名。所有参与者在初始测试(前测)后共接受6次复测,对应6个测试节点:post1至post6,但本数据集不含post5的数据,因二年级期末未采集口头叙事样本。 数据集分为两个子语料库,分别为俄-德双语叙事语料库与土-德双语叙事语料库: 一、俄-德双语儿童叙事语料库 该语料库包含两个文件夹,分别对应俄语与德语叙事,涵盖5个测试节点。本部分总计转写文本100份,涉及儿童10名,整体年龄范围为2;9-10;4岁。 1. 俄语叙事各测试点儿童年龄范围: post1:2;9-4;3(幼儿园阶段) post2:3;9-5;2(幼儿园阶段) post3:4;9-6;1(幼儿园阶段) post4:6;9-7;6(小学一年级期末) post6:8;7-9;10(小学三年级期末) 2. 德语叙事各测试点儿童年龄范围: post1:2;10-4;3(幼儿园阶段) post2:3;9-5;3(幼儿园阶段) post3:4;9-6;2(幼儿园阶段) post4:6;9-7;6(小学一年级期末) post6:8;8-10;4(小学三年级期末) 二、土-德双语儿童叙事语料库 该语料库包含两个文件夹: 1. 第一个文件夹包含前3个测试节点采集的德语叙事,可用于分析单一语言下的早期叙事发展。本部分总计转写文本30份,涉及儿童10名,整体年龄范围为3;5-6;4岁。其德语叙事各测试点儿童年龄范围: post1:3;5-4;3(幼儿园阶段) post2:4;4-5;4(幼儿园阶段) post3:5;3-6;4(幼儿园阶段) 2. 第二个文件夹包含自post2起的4个测试节点采集的土耳其语与德语叙事,可用于分析两种语言直至小学三年级的叙事发展。本部分总计转写文本80份,涉及儿童10名,整体年龄范围为3;10-9;9岁。 - 土耳其语叙事各测试点儿童年龄范围: post2:3;10-5;1(幼儿园阶段) post3:4;9-6;1(幼儿园阶段) post4:6;5-7;8(小学一年级期末) post6:8;6-9;9(小学三年级期末) - 德语叙事各测试点儿童年龄范围: post2:4;1-5;4(幼儿园阶段) post3:5;1-6;4(幼儿园阶段) post4:6;6-7;8(小学一年级期末) post6:8;5-9;8(小学三年级期末) 数据集文件遵循以下命名范式:儿童编码(字母代表儿童第一语言:r代表俄语,t代表土耳其语)、测试工具(MAIN)、故事素材(bb对应《幼鸟》(Baby Birds),bg对应《幼羊》(Baby Goats))、采集语言(de/ru/tr)、测试节点(1=post1,2=post2等)、儿童年龄(年/月)。示例文件名:r009_MAIN_bb_de_4_610。



