遇见数据集

hausa_voa_topics

收藏
Opencsg2024-07-19 更新2025-05-03 收录
官方服务:

资源简介:

Hausa VOA新闻主题分类数据集(hausa_voa_topics)专注于豪萨语新闻标题的主题分类任务。它包含从美国之音豪萨语网站收集的新闻标题,并标注了对应的主题标签,如尼日利亚、非洲、世界、健康或政治。数据集规模在1K到10K之间,分为训练集、验证集和测试集,主要用于文本分类任务。该数据集由专家生成标注,但授权许可等详细信息缺失,需要进一步补充。

Hausa VOA News Topic Classification Dataset (hausa_voa_topics) focuses on the topic classification task for Hausa news headlines. It comprises news headlines collected from the Hausa-language website of Voice of America (VOA), with corresponding topic labels including Nigeria, Africa, World, Health, or Politics. The dataset contains between 1,000 and 10,000 samples, and is partitioned into training, validation, and test sets, primarily intended for text classification tasks. The annotations for this dataset were generated by domain experts, but detailed information such as its licensing terms is missing and requires further supplementation.

创建时间:
2024-07-19
搜集汇总
数据集介绍
hausa_voa_topics 数据集图片
背景与挑战
背景概述
该数据集是一个豪萨语新闻标题主题分类数据集,包含从美国之音豪萨语网站收集的新闻标题,标注了尼日利亚、非洲、世界、健康或政治等主题标签。数据集规模在1K到10K之间,划分为训练集、验证集和测试集,专门用于文本分类任务,支持豪萨语自然语言处理研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务