YT-Data/orpheusplus-tts-enja
收藏官方服务:
资源简介:
该数据集是一个包含训练分割的大规模文本数据集,主要用于自然语言处理任务。数据以input_ids的整数列表形式存储,每个input_id为int32类型。训练集包含5,957,779个样本,总大小约为129.78 GB,下载大小约为129.80 GB。数据集未提供具体的任务描述或来源信息,但根据特征命名推测可能用于语言模型预训练或文本生成任务。
This dataset is a large-scale text dataset containing a training split, primarily used for natural language processing tasks. The data is stored in the form of integer lists as input_ids, with each input_id being of int32 type. The training set consists of 5,957,779 samples, with a total size of approximately 129.78 GB and a download size of approximately 129.80 GB. The dataset does not provide specific task descriptions or source information, but based on feature naming, it is inferred to be potentially used for language model pre-training or text generation tasks.
提供机构:
YT-Data


