遇见数据集

DeL-TaiseiOzaki/Tengentoppa-sft-qwen2.5-32b-reasoning-100k

收藏
Hugging Face2024-11-03 更新2024-12-14 收录
官方服务:

资源简介:

该数据集是通过大規模言語モデル(Qwen2.5-32B-instruct)自动生成的日本語指示与响应集合,包含125,000个样本,主要用于日本語指示付与型タスク的学習和評価。数据集采用JSONL格式,每个样本包含指示文、回答手順、初期回答和精査後回答。数据集生成过程中使用了多種ペルソナ和Chain-of-Thought (CoT) 技术以提高生成多样性。

This dataset is a collection of Japanese instructions and responses automatically generated using a large-scale language model (Qwen2.5-32B-instruct), containing 125,000 samples, primarily used for learning and evaluation of Japanese instruction-giving tasks. The dataset is in JSONL format, with each sample containing an instruction, reasoning steps, initial answer, and refined answer. The generation process utilized multiple personas and Chain-of-Thought (CoT) techniques to enhance diversity.

提供机构:
DeL-TaiseiOzaki
二维码
社区交流群
二维码
科研交流群
商业服务