遇见数据集

Aipresso/long_over_1k_tokens_prompts

收藏
Hugging Face2025-10-22 更新2025-10-25 收录
官方服务:

资源简介:

这是一个专门收集的、包含超过1000个token的长篇英语提示的数据集,用于训练需要广泛上下文和复杂推理的高级模型。数据集包含289个提示,每个提示的token数量在1001到10000之间,适合用于长上下文模型的训练、复杂推理和分析任务、文档级理解系统、高级企业AI应用以及扩展上下文建模研究。

This dataset is a specialized collection of long-form English prompts (≥ 1,000 tokens) designed for training advanced models that require extensive context and complex reasoning. It contains 289 prompts with token counts ranging from 1,001 to 10,000, suitable for long-context model training, complex reasoning and analysis tasks, document-level understanding systems, advanced enterprise AI applications, and research on extended-context modeling.

提供机构:
Aipresso
二维码
社区交流群
二维码
科研交流群
商业服务