hausa_voa_topics
收藏资源简介:
Hausa VOA新闻主题分类数据集(hausa_voa_topics)专注于豪萨语新闻标题的主题分类任务。它包含从美国之音豪萨语网站收集的新闻标题,并标注了对应的主题标签,如尼日利亚、非洲、世界、健康或政治。数据集规模在1K到10K之间,分为训练集、验证集和测试集,主要用于文本分类任务。该数据集由专家生成标注,但授权许可等详细信息缺失,需要进一步补充。
Hausa VOA News Topic Classification Dataset (hausa_voa_topics) focuses on the topic classification task for Hausa news headlines. It comprises news headlines collected from the Hausa-language website of Voice of America (VOA), with corresponding topic labels including Nigeria, Africa, World, Health, or Politics. The dataset contains between 1,000 and 10,000 samples, and is partitioned into training, validation, and test sets, primarily intended for text classification tasks. The annotations for this dataset were generated by domain experts, but detailed information such as its licensing terms is missing and requires further supplementation.




