Grayling-Data/ai-training-bundle
收藏资源简介:
Multi-Domain AI Training Dataset Bundle 是一个包含9个结构化数据集的集合,涵盖分类、指令调整和多轮对话格式,适用于微调大型语言模型(LLMs)和训练自然语言处理(NLP)分类器。数据集包括情感分析(999条记录,通用领域,3类分类:正面、负面、中性)、意图检测(800条记录,客户支持领域,5类分类:购买意图、支持请求、投诉、一般查询、取消)、毒性检测(600条记录,内容审核领域,2类分类:有毒/非有毒)、指令数据集(共1600条记录,覆盖客户支持、英国房地产、个人财务和Python编程领域,采用Alpaca格式)以及对话数据集(共400条记录,覆盖客户支持和编程助手领域,采用ShareGPT格式,每对话4轮)。数据以JSONL和CSV格式提供,支持多种微调框架如Axolotl、LlamaFactory、OpenAI Fine-tuning、HuggingFace TRL和Unsloth。数据集每周更新,由Grayling Data提供,适用于研究、评估和商业用途(需联系提供商获取商业许可)。
The Multi-Domain AI Training Dataset Bundle is a collection of 9 structured datasets across classification, instruction-tuning, and multi-turn conversation formats, ready for fine-tuning LLMs and training NLP classifiers. It includes sentiment analysis (999 records, general domain, 3-class classification: positive, negative, neutral), intent detection (800 records, customer support domain, 5-class classification: purchase intent, support request, complaint, general enquiry, cancellation), toxicity detection (600 records, content moderation domain, binary classification: toxic/non-toxic), instruction datasets (total 1600 records, covering customer support, UK real estate, personal finance, and Python coding domains, in Alpaca format), and conversation datasets (total 400 records, covering customer support and coding assistant domains, in ShareGPT format, with 4 turns per conversation). Data is provided in JSONL and CSV formats, compatible with fine-tuning frameworks such as Axolotl, LlamaFactory, OpenAI Fine-tuning, HuggingFace TRL, and Unsloth. The dataset is updated weekly, provided by Grayling Data, and suitable for research, evaluation, and commercial use (contact provider for commercial licensing).



