carosh/cli-1m
收藏资源简介:
CLI-1M是一个行业多样化的自然语言到Shell命令训练语料库,包含975,933个自然语言到shell命令对,覆盖18个行业领域(包括包管理、云计算、数据库、运维、安全等)、6种shell环境(bash、zsh、fish、powershell、nu、oils-osh)和13种语言(英语、中文、德语、西班牙语、法语、日语、意大利语、葡萄牙语、俄语、阿拉伯语、印地语、韩语、希伯来语)。数据集主要用于指令微调(SFT)和直接偏好优化(DPO),支持文本生成和语言建模任务。数据来源包括导入的结构化文档(如brew/asdf插件注册表和tldr-pages)、LLM合成(使用Claude Haiku 4.5)、跨shell复制和LLM翻译。每个数据行都包含许可证信息(license_spdx字段),排除了GPL/LGPL许可证的数据。数据集提供多种配置:default配置包含SFT训练对和验证集;sample配置为50,000行的分层浏览友好子集;dpo配置包含约33k个DPO偏好对;cross_shell配置包含约410k个跨shell变体;domains配置按行业划分数据。数据集还包含质量层级(导入、合成、策划)和详细的多样性统计。
CLI-1M is an industry-diverse natural language to shell command training corpus containing 975,933 natural-language to shell-command pairs across 18 industries (including package management, cloud, database, devops, security, etc.), 6 shells (bash, zsh, fish, powershell, nu, oils-osh) and 13 languages (English, Chinese, German, Spanish, French, Japanese, Italian, Portuguese, Russian, Arabic, Hindi, Korean, Hebrew). The dataset is primarily designed for instruction fine-tuning (SFT) and direct preference optimization (DPO), supporting text-generation and language-modeling tasks. Data sources include imported structured documentation (e.g., brew/asdf plugin registry and tldr-pages), LLM synthesis (using Claude Haiku 4.5), cross-shell replication, and LLM translation. Each row carries a license field (license_spdx) with GPL/LGPL-licensed sources excluded. The dataset offers multiple configurations: the default config includes SFT training pairs and a validation set; the sample config provides a 50,000-row stratified browse-friendly subset; the dpo config contains ~33k DPO preference pairs; the cross_shell config includes ~410k cross-shell variants; and the domains config splits data by industry. The dataset also includes quality tiers (imported, synthesized, curated) and detailed diversity statistics.



