firefly-train-1.1M
收藏官方服务:
资源简介:
收集了23种常见的中文NLP任务的数据,并且构造了许多与中华文化相关的数据,如对联、作诗、文言文翻译、散文、金庸小说等。对于每个任务,由人工书写若干种指令模板,保证数据的高质量与丰富度,数据量为115万
This dataset compiles data for 23 prevalent Chinese natural language processing (NLP) tasks, and additionally constructs multiple culturally relevant Chinese datasets covering scenarios including couplet generation, poetry composition, classical Chinese text translation, prose collections, and Jin Yong's novel-related tasks. For each task, several manually written instruction templates are developed to ensure the high quality and richness of the data, with a total scale of 1.15 million data samples.
创建时间:
2024-03-21
搜集汇总
数据集介绍

背景与挑战
背景概述
该数据集名为firefly-train-1.1M,是一个中文自然语言处理数据集,包含115万条数据,覆盖23种常见任务,并特别融入了中华文化元素如对联和文言文翻译。其特点在于通过人工设计的指令模板保证数据质量和多样性,适用于训练和评估中文NLP模型。
以上内容由遇见数据集搜集并总结生成



