出门问问序列猴子开源数据集
收藏资源简介:
序列猴子是出门问问提供的超大规模语言模型,基于其通用的表示与推理能力,支持多轮交互,能够大幅度提高生产效率和数据处理能力,被广泛应用于问答系统、自然语言处理、机器翻译、文本摘要等领域。 序列猴子数据集是用于训练序列猴子模型的数据集合,现选择部分数据集向公众开放。 序列猴子开源数据集1.0为序列猴子数据集的首个开源版本,涉及以下领域:中文通用文本语料、古诗今译语料、文本生成语料。
Sequence Monkey is a large-scale language model provided by Mobvoi, leveraging its general representation and reasoning capabilities to support multi-turn interactions, significantly enhancing productivity and data processing efficiency. It is widely applied in areas such as question-answering systems, natural language processing, machine translation, and text summarization. The Sequence Monkey dataset is a collection of data used to train the Sequence Monkey model, with a portion of the dataset now being made publicly available. Sequence Monkey Open Dataset 1.0 is the first open-source version of the Sequence Monkey dataset, covering the following domains: general Chinese text corpus, classical poetry translation corpus, and text generation corpus.
出门问问序列猴子开源数据集概述
数据集版本
- 序列猴子开源数据集1.0:首个开源版本。
数据集内容
- 中文通用文本语料
- 古诗今译语料
- 文本生成语料
- AI配音多风格分类音频语料
应用领域
- 问答系统
- 自然语言处理
- 机器翻译
- 文本摘要
使用许可
- Apache 2.0许可协议:允许自由共享和改编,但需遵循不施加附加限制的条款。
更新日志
- 2024-01-31:首次发布。
- 2024-05-10:添加风格分类音频语料。




