遇见数据集

Diachronic Bangla Corpus, cleaned (pre 1950 and 1950–2025)

收藏
Zenodo2026-07-01 更新2026-08-01 收录
官方服务:

资源简介:

Cleaned diachronic Bangla corpus for the paper "Mapping the Semantic Frontier of Bangla (1950–2025): A Hybrid N-gram and Procrustes-Aligned Embedding Analysis of Diachronic Lexical Change". This is the data record; the analysis code is released separately on GitHub and links back to this DOI. Corpus: about 23.8 million cleaned tokens across five eras — a pre-1950 anchor plus four main eras (1950–1970, 1970–1990, 1990–2010, 2010–2025) — built from human-typed digital editions and modern newspaper text rather than OCR (measured character error rate about 0.7%). Files: one text file per era (pre_1950.txt, 1950_1970.txt, 1970_1990.txt, 1990_2010.txt, 2010_2025.txt) plus stats.json. Each file is the corpus after preprocessing: stemmed, stopword-removed, and sentence-segmented (sentence boundaries are marked), with word order preserved within sentences. This is processed text, not verbatim — the original works cannot be reconstructed exactly. Reproducing the paper: the trained embeddings, drift and ChangeScore tables, statistical-law tests, and figures are all regenerated from this corpus by the analysis code (see the linked GitHub repository). Training is deterministic (fixed seed, single worker), so retraining on this corpus reproduces the same results.

提供机构:
Zenodo
创建时间:
2026-07-01
二维码
社区交流群
二维码
科研交流群
商业服务