遇见数据集

makayel/documentation-kubernetes

收藏
Hugging Face2026-04-07 更新2026-04-12 收录
官方服务:

资源简介:

--- size_categories: n<1K tags: - synthetic - datadesigner configs: - config_name: data data_files: data/*.parquet default: true --- <div style="display: flex; justify-content: space-between; align-items: flex-end; width: 100%; margin-bottom: 1rem;"> <h1 style="flex: 1; margin: 0;">Documentation-Kubernetes</h1> <sub style="white-space: nowrap;">Made with ❤️ using 🦥 Unsloth Studio</sub> </div> --- kubernetes documentation was generated with Unsloth Recipe Studio. It contains 99 generated records. --- ## 🚀 Quick Start ```python from datasets import load_dataset # Load the main dataset dataset = load_dataset("makayel/documentation-kubernetes", "data", split="train") df = dataset.to_pandas() ``` --- ## 📊 Dataset Summary - **📈 Records**: 99 - **📋 Columns**: 3 - **✅ Completion**: 99.0% (100 requested) --- ## 📋 Schema & Statistics | Column | Type | Column Type | Unique (%) | Null (%) | Details | |--------|------|-------------|------------|----------|---------| | `llm_structured_1` | `dict` | llm-structured | 99 (100.0%) | 0 (0.0%) | Tokens: 72 out / 390 in | --- ## ⚙️ Generation Details Generated with 3 column configuration(s): - **llm-structured**: 1 column(s) - **seed-dataset**: 2 column(s) 📄 Full configuration available in [`builder_config.json`](builder_config.json) and detailed metadata in [`metadata.json`](metadata.json). --- ## 📚 Citation If you use Data Designer in your work, please cite the project as follows: ```bibtex @misc{nemo-data-designer, author = {The NeMo Data Designer Team, NVIDIA}, title = {NeMo Data Designer: A framework for generating synthetic data from scratch or based on your own seed data}, howpublished = {\url{https://github.com/NVIDIA-NeMo/DataDesigner}}, year = 2026, note = {GitHub Repository}, } ``` --- ## 💡 About NeMo Data Designer NeMo Data Designer is a general framework for generating high-quality synthetic data that goes beyond simple LLM prompting. It provides: - **Diverse data generation** using statistical samplers, LLMs, or existing seed datasets - **Relationship control** between fields with dependency-aware generation - **Quality validation** with built-in Python, SQL, and custom local and remote validators - **LLM-as-a-judge** scoring for quality assessment - **Fast iteration** with preview mode before full-scale generation For more information, visit: [https://github.com/NVIDIA-NeMo/DataDesigner](https://github.com/NVIDIA-NeMo/DataDesigner) (`pip install data-designer`)

样本量范围:n<1000 标签: - 合成数据(synthetic) - 数据设计器(datadesigner) 配置项: - 配置名称:data 数据文件:data/*.parquet 设为默认配置:是 --- <div style="display: flex; justify-content: space-between; align-items: flex-end; width: 100%; margin-bottom: 1rem;"> <h1 style="flex: 1; margin: 0;">文档-Kubernetes</h1> <sub style="white-space: nowrap;">❤️ 使用 🦥 Unsloth Studio 制作</sub> </div> --- 本数据集为使用 Unsloth Recipe Studio 生成的 Kubernetes 文档数据集,共包含99条生成样本。 --- ## 🚀 快速上手 python from datasets import load_dataset # 加载主数据集 dataset = load_dataset("makayel/documentation-kubernetes", "data", split="train") df = dataset.to_pandas() --- ## 📊 数据集概览 - **📈 样本量**:99 - **📋 字段数**:3 - **✅ 完成率**:99.0%(共请求生成100条样本) --- ## 📋 数据架构与统计信息 | 字段名 | 数据类型 | 字段类别 | 唯一值占比 | 空值占比 | 详情 | |--------|------|-------------|------------|----------|---------| | `llm_structured_1` | `dict` | llm-structured | 99 (100.0%) | 0 (0.0%) | 令牌(Token):输出72个 / 输入390个 | --- ## ⚙️ 生成详情 本数据集采用3字段配置生成: - **llm-structured**:1个字段 - **种子数据集(seed-dataset)**:2个字段 📄 完整配置可在 [`builder_config.json`](builder_config.json) 中查看,详细元数据可在 [`metadata.json`](metadata.json) 中获取。 --- ## 📚 引用规范 若您在工作中使用数据设计器(datadesigner),请按以下格式引用该项目: bibtex @misc{nemo-data-designer, author = {The NeMo Data Designer Team, NVIDIA}, title = {NeMo Data Designer: A framework for generating synthetic data from scratch or based on your own seed data}, howpublished = {url{https://github.com/NVIDIA-NeMo/DataDesigner}}, year = 2026, note = {GitHub Repository}, } --- ## 💡 关于 NeMo 数据设计器(NeMo Data Designer) NeMo 数据设计器是一款用于生成高质量合成数据的通用框架,其能力远超简单的大语言模型(LLM/Large Language Model)提示工程。它提供: - **多样化数据生成**:支持使用统计采样器、大语言模型或现有种子数据集进行数据生成 - **字段关系管控**:实现具备依赖感知能力的字段关联生成 - **质量校验**:内置Python、SQL以及自定义本地/远程校验器 - **LLM 即裁判(LLM-as-a-judge)**:用于质量评估的评分机制 - **快速迭代**:支持在全量生成前通过预览模式进行迭代调试 如需了解更多信息,请访问:[https://github.com/NVIDIA-NeMo/DataDesigner](https://github.com/NVIDIA-NeMo/DataDesigner)(可通过 `pip install data-designer` 安装)

提供机构:
makayel
二维码
社区交流群
二维码
科研交流群
商业服务