遇见数据集

Supplementary materials for: An Audited Bilingual Corpus and Annotation Protocol for the Mother Archetype in Russian and Uzbek Women's Prose

收藏
Zenodo2026-07-19 更新2026-08-02 收录
官方服务:

资源简介:

This deposit accompanies the manuscript "An Audited Bilingual Corpus and Annotation Protocol for the Mother Archetype in Russian and Uzbek Women's Prose" (D. Murodova, G. Salimova), submitted to UBMK 2026. The corpus comprises 39,401 word tokens and 4,871 sentences drawn from seven works — four Russian texts by Masha Traub and three Uzbek texts by Zulfiya Qurolboy qizi. Every work is documented individually by edition, source, coverage, script, and analysed length. A corpus audit identifies three properties that constrain inference: a 1.61:1 language imbalance by tokens, asymmetric source completeness (Russian excerpts vs. complete Uzbek stories), and work-level concentration (one work holds approximately 49% of Russian tokens, one approximately 54% of Uzbek tokens). A five-category multi-label annotation scheme is defined for maternal sub-roles (Mother-Protector, Mother-Educator, Mother-Sacrifice, Mother-Keeper of Tradition, Mother-Identity Seeker). A pilot annotation of 80 balanced sentences by two independent annotators yielded Cohen's κ = 0.662 and Krippendorff's α = 0.668 (substantial agreement). Full corpus-wide annotation is in progress; no distributional estimates are reported in this version. This version contains the corpus manifest, annotation guidelines, annotation schema, and code for candidate extraction, inter-annotator agreement, TF-IDF/SVM baseline, and P@10 retrieval evaluation. Sub-role labels will be added once annotation is complete. Source texts are not redistributed; the manifest provides editions and public URLs for reconstruction.

提供机构:
Zenodo
创建时间:
2026-07-16
二维码
社区交流群
二维码
科研交流群
商业服务