cnnnnnutt/my-clean-orchestra-1M
收藏资源简介:
该数据集是一个大规模的中文文本数据集,包含1,037,091个训练样本和54,584个测试样本。特征包括id(唯一标识符)、title(标题)、group_index(分组索引)、type(类型,可能表示文本类别)、dynasty(朝代,可能指历史时期)、author(作者)和content(内容,即文本主体)。基于特征名称,数据集可能涵盖古典文学、历史文档或其他中文文本,如诗歌、文章等,用于自然语言处理任务,如文本分类、生成或分析。数据集总大小约为301MB,下载大小约为243MB。
This dataset is a large-scale Chinese text dataset comprising 1,037,091 training examples and 54,584 test examples. Features include id (unique identifier), title (title), group_index (group index), type (type, possibly indicating text category), dynasty (dynasty, likely referring to historical period), author (author), and content (content, the main text body). Based on feature names, the dataset may cover classical literature, historical documents, or other Chinese texts, such as poems or articles, for natural language processing tasks like text classification, generation, or analysis. The total dataset size is approximately 301MB, with a download size of about 243MB.



