Annotated Latvian Borrowing Dataset
收藏资源简介:
This dataset contains 3,130 manually annotated samples for the task of lexical borrowing detection in Latvian, created as part of the paper "A Joint Detection Framework for Latvian Loanwords and CalquesUsing Monolingual Data" accepted at LREC 2026. It is the first linguistic resource of its kind, specifically designed to distinguish between Local word (native vocabulary), Loanword (phonological/orthographic borrowings), and Calque (structural borrowings).All samples were extracted from the Latvian Wikipedia corpus (dump dated 02-Feb-2025). The annotation was performed manually by the authors under guidelines developed in consultation with Latvian linguistics experts and native speakers. The classification criteria were validated against multilingual dictionaries (e.g., letonika.lv, termini.gov.lv) and traditional text corpora (dainuskapis).The classification scheme pragmatically extends the traditional definition of Calque to include hybrid compounds, optimizing the data for computational modeling. The dataset intentionally preserves the natural, unbalanced distribution of categories as found in the source corpus to ensure the model learns from a realistic linguistic environment.The robustness of the annotation scheme was confirmed by an inter-annotator agreement (IAA) study, achieving an average Cohen's Kappa (к) coefficient exceeding 0.83.This resource is intended for the training and evaluation of supervised models for lexical borrowing detection in Latvian.



