jchang5/ag_news_annotated
收藏资源简介:
ag_news_annotated数据集是一个基于AG News新闻数据集的人工标注版本,用于文本分类和实体识别任务。数据集包含新闻文本,每个实例都有文本字段,并标注了分类标签(包括商业、科技、体育和世界新闻)和实体标注。数据集规模在10万到100万条记录之间,仅包含训练集分割。它使用Argilla平台创建,支持通过HuggingFace datasets库加载,适用于自然语言处理研究和模型训练。
The ag_news_annotated dataset is a manually annotated variant based on the original AG News dataset, designed for text classification and named entity recognition tasks. It contains news texts, where each instance has a text field, accompanied by annotated classification labels (including Business, Technology, Sports, and World News) and entity annotations. The dataset has a size ranging from 100,000 to 1,000,000 records and only includes the training split. It was constructed using the Argilla platform, supports loading via the Hugging Face Datasets library, and is suitable for natural language processing research and model training.



