遇见数据集

Shrishti-S/ag_news_annotated

收藏
Hugging Face2026-05-18 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个使用Argilla创建的新闻文本数据集,包含新闻文章文本,用于文本分类和实体标注任务。数据集包含一个训练集拆分,大小在10万到100万条记录之间。字段包括文本内容,问题涉及将文本分类为Business、Sci/Tech、Sports、World四个类别,并标注文本中的实体。数据集基于AG News数据集进行标注,但具体来源和标注细节未在README中详细说明。

This dataset is a news text dataset created using Argilla, containing news article texts for text classification and entity annotation tasks. The dataset includes a single train split, with a size between 100K and 1M records. Fields consist of text content, and questions involve classifying text into four categories: Business, Sci/Tech, Sports, World, as well as annotating entities in the text. The dataset is annotated based on the AG News dataset, but specific source and annotation details are not elaborated in the README.

提供机构:
Shrishti-S
二维码
社区交流群
二维码
科研交流群
商业服务