The textual statistics for DialectCorpus.
收藏资源简介:
In this study, we present the acquisition and categorization of a geographically-informed, multi-dialectal Albanian National Corpus, derived from Twitter data. The primary dialects from three distinct regions—Albania, Kosovo, and North Macedonia—are considered. The assembled publicly available dataset encompasses anonymized user information, user-generated tweets, auxiliary tweet-related data, and annotations corresponding to dialect categories. Utilizing a highly automated scraping approach, we initially identified over 1,000 Twitter users with discernible locations who actively employ at least one of the targeted Albanian dialects. Subsequent data extraction phases yielded an augmentation of the preliminary dataset with an additional 1,500 Twitterers. The study also explores the application of advanced geotagging techniques to expedite corpus generation. Alongside experimentation with diverse classification methodologies, comprehensive feature engineering and feature selection investigations were conducted. A subjective assessment is conducted using human annotators, which demonstrates that humans achieve significantly lower accuracy rates in comparison to machine learning (ML) models. Our findings indicate that machine learning algorithms are proficient in accurately differentiating various Albanian dialects, even when analyzing individual tweets. A meticulous evaluation of the most salient attributes of top-performing algorithms provides insights into the decision-making mechanisms utilized by these models. Remarkably, our investigation revealed numerous dialectal patterns that, despite being familiar to human annotators, have not been widely acknowledged within the broader scientific community.
本研究构建并标注了一套基于地理信息的多方言阿尔巴尼亚语国家语料库,该语料库源自Twitter平台数据。本研究覆盖阿尔巴尼亚、科索沃及北马其顿三个不同区域的主流阿尔巴尼亚方言。最终构建的公开可用数据集包含匿名化用户信息、用户发布的推文、推文相关辅助数据,以及对应方言类别的标注标签。本研究首先通过高度自动化的爬虫采集方法,筛选出1000余名拥有明确地理位置、且至少使用一种目标阿尔巴尼亚方言的活跃推特用户;后续的数据采集阶段又新增1500名推特用户,扩充了初始数据集。本研究同时探索了先进地理标记技术的应用,以加速语料库的构建流程。本研究针对多种分类方法开展对比实验,并完成了全面的特征工程与特征选择研究。本研究通过人工标注者开展主观评估实验,结果显示,人工标注的准确率显著低于机器学习(Machine Learning,ML)模型。研究结果表明,即便仅针对单条推文进行分析,机器学习算法也能够精准区分不同类别的阿尔巴尼亚方言。通过对性能最优算法的核心显著特征进行细致评估,本研究揭示了此类模型的决策机制。值得注意的是,本研究发现了多种方言特征模式——这些模式虽为人工标注者所熟知,但尚未在全球学术界得到广泛认可。



