遇见数据集

Long Data Collections

收藏
arXiv2025-09-30 收录
官方服务:

资源简介:

该数据集是一组用于在不超过32k个令牌的序列上训练模型的长期上下文数据集的集合。此外,部分数据集也被用于提高针对长上下文的指令调整效果。所涉及的任务是指令调整。

This dataset collection comprises a set of long-context datasets tailored for training models on sequences with no more than 32k tokens. Additionally, a portion of these datasets is utilized to improve the performance of long-context instruction tuning. The core task involved here is instruction tuning.

搜集汇总
数据集介绍
Long Data Collections 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务