utter-project/EuroBlocks-SFT-2512
收藏官方服务:
资源简介:
该数据集包含对话内容,每条记录包含用户和助手的消息列表(conversations)以及对话使用的语言(language)。语言标注可能不完全准确,尤其是对话中包含多种语言时。数据集分为训练集,包含1,094,265个例子,总大小为4,094,880,720字节。
The dataset contains conversational data, with each record including a list of messages from both user and assistant (conversations) and the language used in the conversation (language). The language annotation may not be fully accurate, especially when the conversation includes multiple languages. The dataset is split into a training set with 1,094,265 examples and a total size of 4,094,880,720 bytes.
提供机构:
utter-project


