遇见数据集

Bartangi Lemmatization & Embedding Resource (v1.0)

收藏
Zenodo2025-12-06 更新2026-05-26 收录
官方服务:

资源简介:

Bartangi Lemmatization & Embedding Resource (v1.0) This release provides a reproducible Bartangi NLP resource: a rule-based lemmatizer (exceptions + ordered suffix rules), a normalized lemma corpus, Word2Vec embeddings (CBOW/Skip-gram), evaluation scripts, figures, and a 300-token gold review set for intrinsic scoring. What’s inside- rules/ — exceptions lexicon and ordered suffix rules (documented).- artifacts/lemmas.txt — normalized, lemmatized corpus (derived artifact).- models/ — w2v_cbow.model, w2v_sg.model (Gensim).- results/ — downstream & ablation CSVs; intrinsic evaluation.- figures/ — PCA/t-SNE plots.- gold/ — lemma_gold_review.csv (300 tokens) + scoring script.- config.yaml, requirements.txt, versions.json, README.md. Key results (v1.0)- Intrinsic lemmatizer accuracy: 99.33% on a 300-token review set (53 changed tokens).- Downstream (3 seeds): CBOW F1=0.5361, Acc=0.5389; SG F1=0.5206, Acc=0.5222.- Ablation (CBOW): Raw 0.5150/0.5167 → Lemma 0.5361/0.5389 (+0.021 F1, +0.022 Acc). Ethics & licensingWe redistribute only derived artifacts (lemmas, rules, embeddings, models) and do not republish raw source texts. A provenance table with source links and licenses is included (artifacts/stats/provenance.csv). Personal/sensitive items were filtered during preprocessing. As Bartangi is an endangered language, these resources are shared to support documentation and research. ReproducibilityEnvironment snapshot in requirements.txt and versions.json. Seeds: 42/43/44; Word2Vec: 200d, window=5, epochs=20. See README for run instructions.

提供机构:
Zenodo
创建时间:
2025-12-06
二维码
社区交流群
二维码
科研交流群
商业服务