pfm1-dataset
收藏资源简介:
该数据集是一个中文新闻文本分类数据集,专门设计用于研究和评估文本分类模型,特别是针对新闻文本的细粒度分类任务。数据集包含1,000条中文新闻文本样本,按照7:1:2的比例划分为训练集、验证集和测试集。每条样本包含以下字段:唯一标识符id、新闻标题title、新闻正文content、数字标签label(0-7)以及对应的中文类别标签label_zh。数据集涵盖8个新闻类别,采用多分类标签体系。数据来源于公开的新闻数据,适用于文本分类模型的训练、验证和测试。
This dataset is a Chinese news text classification dataset, specifically designed for researching and evaluating text classification models, particularly for fine-grained classification tasks on news texts. It contains 1,000 Chinese news text samples, divided into training, validation, and test sets in a 7:1:2 ratio. Each sample includes the following fields: unique identifier id, news title title, news content content, numeric label label (0-7), and corresponding Chinese category label label_zh. The dataset covers 8 news categories and uses a multi-class labeling system. The data is sourced from publicly available news data and is suitable for training, validation, and testing of text classification models.




