TF-IDF Matrix for Grades 5–11 Uzbek Textbooks"
收藏资源简介:
This dataset contains the TF-IDF feature matrix extracted from a cleaned corpus of Uzbek school textbooks for grades 5 through 11. Each row corresponds to a document, and each column represents a unique word feature used for thematic text classification and linguistic analysis. The corpus includes 96 textbooks covering various subjects such as literature, history, physics, chemistry, biology, and more. The matrix includes 221,036 unique word features, each transformed into numerical vectors using the Term Frequency–Inverse Document Frequency (TF-IDF) method. These vectors represent the importance of each word in a given document relative to the entire corpus, making the dataset suitable for educational text mining, natural language processing tasks, and corpus linguistics research involving Uzbek-language instructional material.



