Computer-assisted study of historical Lemkian (Transcarpathian Ukraine) lects: basic vocabulary approach - Supplementary material 1 (Dataset)
收藏资源简介:
The dataset presents the historical Lemkian (small territorial lect in Transcarpathian Ukraine) text LA1407 from the earlier non-digitised study (Nakonetschna and Rudnyćkyj 1940). The dataset is in .conllu format with silver morhological tagging and lemmatisation performed by Stanza model for Ukrainian (Qi et al. 2018; Qi et al. 2020). The dataset is also enriched with basic vocabulary and named entities tags. The texts in the dataset are present in three forms: IPA transcription, standard-like Cyrillic transcription, and original German translation. The basic form of representation is standard-like Cyrillic transcription to make automatic tagging easier; however, each token in miscellanea section also has keys wf (normalised form) and tf (transcription in IPA).



