遇见数据集

utter-project/EuroBlocks-SFT-2512

收藏
Hugging Face2026-02-06 更新2026-02-07 收录
官方服务:

资源简介:

该数据集包含对话内容,每条记录包含用户和助手的消息列表(conversations)以及对话使用的语言(language)。语言标注可能不完全准确,尤其是对话中包含多种语言时。数据集分为训练集,包含1,094,265个例子,总大小为4,094,880,720字节。

The dataset contains conversational data, with each record including a list of messages from both user and assistant (conversations) and the language used in the conversation (language). The language annotation may not be fully accurate, especially when the conversation includes multiple languages. The dataset is split into a training set with 1,094,265 examples and a total size of 4,094,880,720 bytes.

提供机构:
utter-project
二维码
社区交流群
二维码
科研交流群
商业服务