M0N0S0DIUM/ZMK_DOCS
收藏资源简介:
ZMK Data Clean 1 是一个合成数据集,使用 Unsloth Recipe Studio 生成,包含 2,500 条记录。该数据集是 NVIDIA NeMo Data Designer 框架的一部分,旨在生成高质量合成数据,支持多样化数据生成、字段间关系控制、质量验证(包括 Python、SQL 和自定义验证器)以及 LLM 作为评判的评分功能。数据集具有 3 列,其中一列为 llm_structured 类型,表示以列表字典形式存储的结构化数据,适用于 NLP 和机器学习任务。
ZMK Data Clean 1 is a synthetic dataset generated with Unsloth Recipe Studio, containing 2,500 records. It is part of the NVIDIA NeMo Data Designer framework, which is designed for generating high-quality synthetic data with features such as diverse data generation using statistical samplers or LLMs, relationship control between fields, quality validation with built-in Python, SQL, and custom validators, LLM-as-a-judge scoring, and fast iteration via preview mode. The dataset has 3 columns, including one llm-structured column in list[dict] format, making it suitable for NLP and machine learning applications.




