STAGES OF FORMING AND DIGITIZING THE AUTHOR CORPUS
收藏资源简介:
This study examines the stages of forming and digitizing the Nusratulla Jumaxo‘ja author corpus within the framework of modern corpus linguistics and digital humanities. The research focuses on the scientific and methodological principles of corpus creation, including source collection, metadata development, text normalization, tokenization, indexing, concordance generation, and statistical analysis. Particular attention is paid to the challenges of digitizing Uzbek-language texts, especially issues related to OCR accuracy, Unicode standardization, and apostrophe encoding in Uzbek Latin script. The study demonstrates that the author corpus is not merely an electronic archive, but a multilayered linguistic platform designed for linguostatistical, stylistic, and semantic analysis. Through concordance and frequency-based analysis, the corpus enables the identification of the author’s idiolect, dominant lexical units, and discursive strategies. The integration of metadata and etymological modules further enhances the analytical capabilities of the system. The research concludes that the Nusratulla Jumaxo‘ja author corpus serves as an important digital resource for Uzbek linguistics, stylometry, lexicography, and corpus-based literary studies, while also offering a methodological model for the development of future author corpora in Uzbek corpus linguistics.
本研究在现代语料库语言学与数字人文的框架下,探讨了努斯拉图拉·朱马霍贾(Nusratulla Jumaxo‘ja)作者语料库的构建与数字化阶段。本研究聚焦于语料库构建的科学与方法论原则,涵盖源数据采集、元数据(Metadata)编制、文本归一化、Token分词、索引构建、语境共现语料生成以及统计分析等环节。本研究特别关注乌兹别克语文本数字化过程中面临的挑战,尤其是与光学字符识别(Optical Character Recognition, OCR)准确率、统一码(Unicode)标准化以及乌兹别克语拉丁字母撇号编码相关的问题。本研究表明,该作者语料库并非单纯的电子档案,而是一个面向语言统计、文体与语义分析的多层级语言学平台。通过语境共现与基于词频的分析,该语料库可用于识别作者的个人语域、核心词汇单元与话语策略。元数据与词源模块的集成进一步提升了该系统的分析能力。本研究最终得出结论:努斯拉图拉·朱马霍贾作者语料库可作为乌兹别克语言学、文体计量学、词典编纂学以及基于语料库的文学研究的重要数字资源,同时也为乌兹别克语料库语言学领域未来作者语料库的构建提供了方法论范式。



