Python_GOD_Coder_Omniforge_AI_12k
收藏资源简介:
# Python GOD Coder Omniforge AI 12k **Creator:** Within Us AI A **12,000-row mixed-format Python coding dataset** designed as a sharpening corpus for building a small but dangerous Python specialist. This dataset is intentionally focused on the practical behaviors that matter for a modern Python coding model: - implementation with tests - strict code-only instruction following - debugging and repair - refactoring for readability and production readiness - next-token code completion - fill-in-the-middle (PSM and SPM) - repository-context completion - code critique and ranking - modern AI Python stack tasks such as FastAPI, vLLM, LangGraph, MCP, PyTorch, asyncio, tool registries, and general production Python utilities ## Splits - **train**: 11760 - **validation**: 240 ## Row distribution ```json { "implement": 2400, "implement_strict": 1200, "debug": 1500, "refactor": 1200, "completion": 1800, "fim_psm": 1200, "fim_spm": 900, "repo_completion": 780, "critique": 420, "test_first": 600 } ``` ## Row families This dataset intentionally mixes several schemas in one corpus. ### 1. Instruction / repair / refactor rows Common keys: - `row_id` - `task_type` - `difficulty` - `skills` - `style_tags` - `instruction` - `input` - `output` - `tests` - `source_template` - `domain` ### 2. Completion rows Common keys: - `row_id` - `task_type` - `difficulty` - `skills` - `style_tags` - `prefix` - `completion` - `tests` - `source_template` - `domain` ### 3. Fill-in-the-middle rows Common keys: - `row_id` - `task_type` - `difficulty` - `skills` - `style_tags` - `fim_mode` - `prefix` - `suffix` - `middle` - `tests` - `source_template` - `domain` ### 4. Repo-context rows Common keys: - `row_id` - `task_type` - `difficulty` - `skills` - `style_tags` - `instruction` - `context_files` - `target_file_path` - `target_file_prefix` - `target_file_suffix` - `answer` - `tests` - `source_template` - `domain` ### 5. Critique rows Common keys: - `row_id` - `task_type` - `difficulty` - `skills` - `style_tags` - `instruction` - `candidate_a` - `candidate_b` - `preferred` - `reason` - `output` - `tests` - `source_template` - `domain` ## Intended use This dataset is meant as a **finishing-tune and sharpening dataset**, especially for a model that already has some general code ability. Recommended uses: - supervised fine-tuning - code completion tuning - FIM tuning - repair / refactor tuning - repo-context tuning - code-review preference expansion ## Important note This is a **synthetic / templated training dataset**, not a public benchmark. It is designed to teach modes of behavior, not to act as a leaderboard by itself. Use separate held-out evaluation sets and private test suites for honest measurement. ## Example loading ```python from datasets import load_dataset ds = load_dataset("json", data_files={ "train": "train.jsonl", "validation": "validation.jsonl", }) print(ds) print(ds["train"][0]) ``` ## Suggested training strategy A strong training recipe for a small Python specialist: 1. start from a code-capable base model 2. fine-tune on your broad Python corpus 3. mix in this dataset as a sharpening pass 4. oversample FIM, repo-context, and debug rows in a short second pass 5. merge the final adapter into the base model if you want a standalone release ## License `other` This dataset is released under the Within Us AI Custom Dataset License v1.0. Include the LICENSE.txt file with any redistribution of the dataset repository.



