遇见数据集

Annotated Latvian Borrowing Dataset

收藏
Zenodo2026-03-13 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains 3,130 manually annotated samples for the task of lexical borrowing detection in Latvian, created as part of the paper "A Joint Detection Framework for Latvian Loanwords and CalquesUsing Monolingual Data" accepted at LREC 2026. It is the first linguistic resource of its kind, specifically designed to distinguish between Local word (native vocabulary), Loanword (phonological/orthographic borrowings), and Calque (structural borrowings).All samples were extracted from the Latvian Wikipedia corpus (dump dated 02-Feb-2025). The annotation was performed manually by the authors under guidelines developed in consultation with Latvian linguistics experts and native speakers. The classification criteria were validated against multilingual dictionaries (e.g., letonika.lv, termini.gov.lv) and traditional text corpora (dainuskapis).The classification scheme pragmatically extends the traditional definition of Calque to include hybrid compounds, optimizing the data for computational modeling. The dataset intentionally preserves the natural, unbalanced distribution of categories as found in the source corpus to ensure the model learns from a realistic linguistic environment.The robustness of the annotation scheme was confirmed by an inter-annotator agreement (IAA) study, achieving an average Cohen's Kappa (к) coefficient exceeding 0.83.This resource is intended for the training and evaluation of supervised models for lexical borrowing detection in Latvian.

提供机构:
Zenodo
创建时间:
2025-10-17
二维码
社区交流群
二维码
科研交流群
商业服务