nickoo004/queryshield-multilingual
收藏资源简介:
QueryShield是一个高质量的多语言提示优化数据集,旨在训练大型语言模型(LLMs)成为30个专业领域的专家级响应者。数据集包含原始用户问题和详细的指令提示,指导下游LLM如何回答问题,而非直接提供答案。数据集涵盖乌兹别克语、卡拉卡尔帕克语、哈萨克语、俄语和英语,包括用户用一种语言提问但要求用另一种语言回答的跨语言场景。数据集总行数约19,530,其中约28%为跨语言对。数据格式为JSONL,每个JSON对象包含用户问题、优化提示、输入输出语言代码、目标角色等字段。数据集适用于指令调整、多语言提示优化和中亚语言支持。
QueryShield is a high-quality synthetic dataset of prompt optimization pairs designed to train LLMs to act as expert-level responders across 30 professional domains. Each row contains a raw user question and a detailed instruction prompt telling a downstream LLM how to answer it — not the answer itself. The dataset is multilingual, covering Uzbek, Karakalpak, Kazakh, Russian, and English, including cross-lingual scenarios where the user writes in one language but requests a response in another. It contains ~19,530 rows, with ~28% being cross-lingual pairs. The data is in JSONL format, with each JSON object including fields like user_question, optimized_prompt, input/output language codes, and target_role. It is intended for instruction tuning, multilingual prompt optimization, and Central Asian language support.




