German Innsbruck Corpus (GermInnC) 1800-1950
收藏资源简介:
<strong>A digital corpus on variation in German (1800-1950)</strong> The <em>German Innsbruck Corpus</em> <em>(GermInnC) 1800-1950</em> is a digitised corpus built after the fashion of the <em>German Manchester Corpus (GerManC) 1650-1800</em> (cf. Scheible et al. 2011; Durrell et al. 2012). Hence, the corpus design of the GermInnC is balanced according to period, region and genre. The GermInnC consists of ca. 840,000 tokens, ca. 120,000 per genre (seven in total: Drama, Humanities, Legal texts, Narrative prose, Newspapers, Scientific texts, Sermons). It is subdivided into three periods, 1800-1850, 1851-1900 und 1901-1950, as well as five regions, North German, West Central German, East Central German, West Upper German (including Switzerland), East Upper German (including Austria). The corpus can be retrieved in a raw version, a lemmatised, fully-annotated version, or an “all data” file (including metadata annotation of file names and periods) for further import and processing. The <em>Stuttgart Tag Set </em>(STTS) and the POS-Tagger <em>TreeTagger </em>was used for linguistic annotation. Two documentation files (word and excel, both included in the download package), provide a more detailed description of the corpus and the digitisation. The corpus may be of interest to all scholars working on the history of the German language, standardisation of German, variation and change, historical sociolinguistics, and Germanic linguistics. The corpus was generously funded by the early career funding of the University of Innsbruck (October 2018 through September 2019).



