遇见数据集

TF-IDF Matrix for Grades 5–11 Uzbek Textbooks"

收藏
Zenodo2025-06-20 更新2026-05-26 收录
官方服务:

资源简介:

This dataset contains the TF-IDF feature matrix extracted from a cleaned corpus of Uzbek school textbooks for grades 5 through 11. Each row corresponds to a document, and each column represents a unique word feature used for thematic text classification and linguistic analysis. The corpus includes 96 textbooks covering various subjects such as literature, history, physics, chemistry, biology, and more. The matrix includes 221,036 unique word features, each transformed into numerical vectors using the Term Frequency–Inverse Document Frequency (TF-IDF) method. These vectors represent the importance of each word in a given document relative to the entire corpus, making the dataset suitable for educational text mining, natural language processing tasks, and corpus linguistics research involving Uzbek-language instructional material.

提供机构:
Zenodo
创建时间:
2025-06-20
二维码
社区交流群
二维码
科研交流群
商业服务