遇见数据集

AnujPatnaik1/ag_news_annotated1

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个新闻文本分类数据集,包含文本(text)和标签(label)两个特征。标签分为四个类别:World(世界新闻)、Sports(体育新闻)、Business(商业新闻)和Sci/Tech(科技新闻)。数据集分为训练集(train)和测试集(test),训练集有120,000个示例,测试集有7,600个示例,总大小约为31.7 MB。数据集用于自然语言处理任务,如文本分类模型训练和评估。

This dataset is a news text classification dataset that includes two features: text and label. The labels fall into four categories: World (World News), Sports (Sports News), Business (Business News), and Sci/Tech (Science and Technology News). The dataset is split into a training set (train) and a test set (test), containing 120,000 samples in the training set and 7,600 samples in the test set, with a total size of approximately 31.7 MB. This dataset is designed for natural language processing (NLP) tasks such as training and evaluating text classification models.

提供机构:
AnujPatnaik1
二维码
社区交流群
二维码
科研交流群
商业服务