common_long_32k
收藏资源简介:
Common-Long-32K是一个正在进行中的数据集,旨在为Comma模型的长上下文训练提供数据支持。该数据集源自Common Pile v0.1的已过滤数据,并进一步通过序列长度进行筛选,保留字符长度在16,000到145,000之间的文本序列。设计基于每4到6个字符对应一个token的假设,目标是构建一个包含约4,000到32,000个token的序列集合(实际范围可能低至2,500到24,000 token)。该数据集计划与LongLoRA技术结合使用,用于训练或微调能够处理长上下文的语言模型。未来还计划推出一个64K版本。
Common-Long-32K is an ongoing dataset designed to provide data support for long-context training of the Comma model. It is derived from filtered data of Common Pile v0.1 and further screened by sequence length, retaining text sequences with character lengths between 16,000 and 145,000. The design is based on the assumption that every 4 to 6 characters correspond to one token, with the goal of constructing a collection of sequences containing approximately 4,000 to 32,000 tokens (actual range may be as low as 2,500 to 24,000 tokens). The dataset is intended to be used in conjunction with LongLoRA technology for training or fine-tuning language models capable of handling long contexts. A 64K version is also planned for the future.
数据集名称
Common-Long-32K
数据集描述
- 状态:工作正在进行中。
- 来源:基于 Common Pile v0.1 已过滤数据的大小过滤版本,用于 Comma 模型的长上下文训练。
- 序列长度:保留了长度为 16k 到 145k 字符的序列(假设每 token 对应 4-6 个字符)。
- 目标上下文长度:旨在覆盖 4k 到 32k token 范围,实际范围可能低至 2.5k 到 24k token。
- 应用场景:将用于 LongLoRA 训练。
- 未来计划:后续会制作 64K 版本。
- 原始数据集合:参考 Common Pile v0.1 过滤数据集合:https://huggingface.co/collections/common-pile/common-pile-v01-filtered-data
其他信息
- 整理者:未提供
- 资金方:未提供
- 共享者:未提供
- 语言:未提供
- 许可证:未提供




