遇见数据集

yibotongxue/Data-Prep-Bench

收藏
Hugging Face2026-05-02 更新2026-05-31 收录
官方服务:

资源简介:

Data-Prep-Bench是一个用于大型语言模型(LLM)监督微调(SFT)和评估的综合数据集。它覆盖六个核心领域:金融、医学、法律、数学、科学和通用领域。该数据集的关键特色是采用了12种不同的数据生成方法,包括基于代理的方法(使用Qwen3.5-Plus、GLM-4.7、Claude Opus 4.6、Gemini 3.0 Pro、GPT-5.2、GPT-5.3-codex等模型)、DataFlow系列(DataFlow和DataFlow Agent)、纯LLM生成方法(使用Claude Opus 4.6、Gemini 3.0 Pro、GPT-5.2)以及SKILL方法(使用Claude Opus 4.6),这些方法处理原始语料库(包括PDF电子书和Markdown文件)以生成高质量的问答对。此外,数据集还提供了标准化的评估基准,涵盖商业、法律和医学三个领域,包含统一的样本结构和元数据。数据集支持多语言(训练语料含中英文,基准为英文),主要用于模型微调、性能评估和数据生成方法研究。

Data-Prep-Bench is a comprehensive dataset designed for supervised fine-tuning (SFT) and evaluation of large language models (LLMs). It spans six core domains: finance, medicine, law, mathematics, science, and general. A key feature of this dataset is the use of 12 distinct data generation methods, including agent-based approaches (using models such as Qwen3.5-Plus, GLM-4.7, Claude Opus 4.6, Gemini 3.0 Pro, GPT-5.2, GPT-5.3-codex), DataFlow series (DataFlow and DataFlow Agent), pure LLM-based generation (using Claude Opus 4.6, Gemini 3.0 Pro, GPT-5.2), and a SKILL method (using Claude Opus 4.6). These methods process raw corpora (including PDF e-books and Markdown files) to produce high-quality question-answer pairs. Additionally, the dataset provides standardized evaluation benchmarks covering business, law, and medicine domains, with a unified sample structure and metadata. The dataset is multilingual (training corpora contain both Chinese and English; benchmarks are in English) and is intended for model fine-tuning, performance evaluation, and research on data generation methodologies.

提供机构:
yibotongxue
二维码
社区交流群
二维码
科研交流群
商业服务