"BUILDING AND OPTIMIZING UZBEK–ENGLISH TERMINOLOGICAL DATABASES USING DATA SCIENCE METHODS"
收藏资源简介:
Abstract. This article examines the construction of Uzbek–English terminological databases and the optimization of existing ones through Data Science methods. The aim of the study is to develop a model for a digital terminological database that ensures terminological correspondence between the two languages, can be automatically expanded, and is subject to quality control. The study employed corpus linguistics, natural language processing (NLP), automatic text tagging, and machine learning algorithms. A parallel corpus comprising technical, scientific, and socio-political texts was compiled as the data source, from which term candidates were extracted using statistical and neural methods. The results show that a hybrid approach — combining rule-based filtering with machine learning models — significantly improves the accuracy of term identification and reduces the time required for manual database population. The structure of the resulting database, its metadata schema, and the quality assessment criteria are also presented. In conclusion, Data Science tools represent an effective and promising direction for developing terminological resources for the Uzbek language.



