Tibetan text classification is a basic task in Tibetan natural language processing. Based on large-scale pre-trained language model and fine-tuning is the current mainstream text classification model.
The H-Prop dataset contains 28,630 articles created by translating a portion of Proppy Corpus in Hindi. Each article is labeled as either “propagandistic” (positive class) or “non-propagandistic” (neg