marketing-instruct-13k
收藏资源简介:
marketing-instruct-13k是一个专门为营销文案生成模型设计的指令调优数据集,包含约13,000个营销文案示例,覆盖五种核心任务类型:产品描述(约4,800个示例,基于属性)、广告/社交文案(约2,800个示例,包括Instagram广告和品牌声音配对)、电子邮件营销(约2,400个示例,涵盖活动邮件和个性化邮件)、品牌声音重写(约1,500个示例,将中性文案重写为特定品牌风格)以及表格到洞察(约1,500个示例,将活动绩效数据总结为洞察报告)。该数据集扩展了早期版本(marketing-instruct-8k),新增了基于属性的产品描述示例,并最终源自首个数据集marketing-instruct-4k。所有示例遵循统一结构,包含instruction、input和output字段。数据来源于10个公开数据集,经过去重、清理等处理,并通过AdaptData平台适配。数据集专用于AutoScientist Challenge 2026(营销类别),已成功微调Marketing-Llama-3.3-70B模型,在基准测试中达到88%胜率。适用场景包括广告文案生成、电子邮件营销等,但排除营销策略等任务。数据集仅含英文内容,基于西方营销惯例,存在偏向产品描述任务、合成示例可能携带生成模型风格等局限性,建议在使用前由人类营销人员审查输出。
marketing-instruct-13k is a curated instruction-tuning dataset for marketing copy generation, containing approximately 13,000 examples across five core task types. It is designed for the AutoScientist Challenge 2026 (marketing category) and is the largest dataset in the Marketing-Mixtral/Llama series, successfully used to fine-tune the Marketing-Llama-3.3-70B model, achieving an 88% win rate in benchmarks. The dataset extends the earlier marketing-instruct-8k version by adding around 4,000 attribute-based product description examples and traces back to the original marketing-instruct-4k dataset. All examples follow a uniform structure with instruction, input, and output fields. Task types include product descriptions (~4,800 examples, attribute-based), ad/social copy (~2,800 examples, including Instagram-specific ads and brand voice pairing), email marketing (~2,400 examples, covering campaign and personalized emails), brand voice rewriting (~1,500 examples, rewriting neutral copy to match a specified brand voice), and table-to-insights (~1,500 examples, summarizing campaign performance data into insights). Data is sourced from 10 public datasets, processed through deduplication, cleaning of short examples and placeholders, and adapted via the AdaptData platform. It is intended for marketing copy generation tasks, excluding strategy and analysis, and is limited to English with Western marketing conventions. Limitations include bias towards product descriptions and potential style patterns from synthetic examples, recommending human review before deployment.




