遇见数据集

German Innsbruck Corpus (GermInnC) 1800-1950

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>A digital corpus on variation in German (1800-1950)</strong> The <em>German Innsbruck Corpus</em> <em>(GermInnC) 1800-1950</em> is a digitised corpus built after the fashion of the <em>German Manchester Corpus (GerManC) 1650-1800</em> (cf. Scheible et al. 2011; Durrell et al. 2012). Hence, the corpus design of the GermInnC is balanced according to period, region and genre. The GermInnC consists of ca. 840,000 tokens, ca. 120,000 per genre (seven in total: Drama, Humanities, Legal texts, Narrative prose, Newspapers, Scientific texts, Sermons). It is subdivided into three periods, 1800-1850, 1851-1900 und 1901-1950, as well as five regions, North German, West Central German, East Central German, West Upper German (including Switzerland), East Upper German (including Austria). The corpus can be retrieved in a raw version, a lemmatised, fully-annotated version, or an “all data” file (including metadata annotation of file names and periods) for further import and processing. The <em>Stuttgart Tag Set </em>(STTS) and the POS-Tagger <em>TreeTagger </em>was used for linguistic annotation. Two documentation files (word and excel, both included in the download package), provide a more detailed description of the corpus and the digitisation. The corpus may be of interest to all scholars working on the history of the German language, standardisation of German, variation and change, historical sociolinguistics, and Germanic linguistics. The corpus was generously funded by the early career funding of the University of Innsbruck (October 2018 through September 2019).

**1800-1950年德语变体数字语料库** 《因斯布鲁克德语语料库(GermInnC)1800-1950》是参照《曼彻斯特德语语料库(GerManC)1650-1800》(参见Scheible等,2011;Durrell等,2012)构建的数字化语料库。该语料库的设计严格遵循时期、区域与文体的均衡配比原则。 GermInnC总计约含84万个Token,涵盖7类文体,每类文体规模约12万个Token,具体包括戏剧、人文类文本、法律文本、叙事散文、报纸、科学文本与布道文。该语料库被划分为三个时期:1800-1850年、1851-1900年及1901-1950年,同时覆盖五大德语区:北德语区、中西德语区、东中德语区、西上德语区(含瑞士)与东上德语区(含奥地利)。 本语料库提供三种获取版本:原始文本版、词形还原且完全标注版,以及「全数据」文件(包含文件名与时期的元数据标注),可用于后续导入与处理。语言标注环节采用了斯图加特标注集(Stuttgart Tag Set, STTS)与词性标注器TreeTagger。 下载包中附带两份文档文件(Word与Excel格式),对本语料库及数字化流程进行了更为详尽的说明。本语料库适用于所有开展德语语言史、德语规范化、语言变体与演变、历史社会语言学及日耳曼语言学研究的学者。 本语料库由因斯布鲁克大学青年学者资助项目(2018年10月至2019年9月)慷慨资助。

提供机构:
Zenodo
创建时间:
2019-09-23
二维码
社区交流群
二维码
科研交流群
商业服务