Shrishti-S/ag_news_annotated
收藏官方服务:
资源简介:
该数据集是一个使用Argilla创建的新闻文本数据集,包含新闻文章文本,用于文本分类和实体标注任务。数据集包含一个训练集拆分,大小在10万到100万条记录之间。字段包括文本内容,问题涉及将文本分类为Business、Sci/Tech、Sports、World四个类别,并标注文本中的实体。数据集基于AG News数据集进行标注,但具体来源和标注细节未在README中详细说明。
This dataset is a news text dataset created using Argilla, containing news article texts for text classification and entity annotation tasks. The dataset includes a single train split, with a size between 100K and 1M records. Fields consist of text content, and questions involve classifying text into four categories: Business, Sci/Tech, Sports, World, as well as annotating entities in the text. The dataset is annotated based on the AG News dataset, but specific source and annotation details are not elaborated in the README.
提供机构:
Shrishti-S


