遇见数据集

abir-hr196/clt_tinystories_tokenized

收藏
Hugging Face2025-07-14 更新2025-10-25 收录
官方服务:

资源简介:

该数据集包含五种语言的文本数据,分别为法语、德语、阿拉伯语、中文和英语。每个语言的数据都包括文本内容和对应的输入ID序列。

The dataset contains text data in five languages: French, German, Arabic, Chinese, and English. Each languages data includes the text content and the corresponding input ID sequences.

提供机构:
abir-hr196
二维码
社区交流群
二维码
科研交流群
商业服务