遇见数据集

Ali-Yaser/CodeX-2M-Qwen36-ChatML

收藏
Hugging Face2026-05-10 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个大规模文本数据集,包含约218.97万个样本,总大小约为58.2 GB。数据特征仅包括一个text字段,存储字符串类型文本。数据集仅提供训练分割(train),数据文件路径为data/train-*。由于README未提供详细描述,无法确定数据集的具体内容、来源或应用领域。

This dataset is a large-scale text dataset containing approximately 2.19 million samples, with a total size of around 58.2 GB. The only data feature is a "text" field that stores string-type text. The dataset only provides the training split (train), and the data file path is data/train-*. Since no detailed description is provided in the README, the specific content, source, or application domain of the dataset cannot be determined.

提供机构:
Ali-Yaser
二维码
社区交流群
二维码
科研交流群
商业服务