pashto-translated-psychology
收藏资源简介:
该数据集包含翻译成普什图语的简短、清晰、领域特定的心理学条目,专为普什图语大语言模型对齐、心理健康推理任务、情感感知对话、短上下文监督微调训练以及低资源语言研究而优化。每个条目长度不超过500字符,确保数据简洁,适用于快速训练、令牌高效的微调和移动友好型模型开发。数据格式为JSON对象,每行包含一个input字段(普什图语提示)和一个output字段(普什图语翻译的心理学回答),例如输入为د اضطراب نښې څه دي?(焦虑的症状是什么?),输出为相应的心理学解释。数据集规模在1,000到10,000个样本之间,属于文本生成和翻译任务范畴,使用Apache 2.0许可证,允许研究和商业用途。预期应用包括构建普什图语指令调优模型、心理健康聊天机器人、领域特定推理系统,并支持对Qwen、LLaMA、Ministral、Gemma等模型的微调。需要注意的是,该数据集不能替代专业心理学建议,翻译可能存在细微风格差异,建议在部署时与安全对齐数据集结合使用以增强模型安全性。
This dataset provides short, clear, domain-specific psychology entries translated into Pashto, optimized for Pashto large language model alignment, mental health reasoning tasks, emotion-aware dialogues, short-context supervised fine-tuning training, and low-resource language research. Each entry is no longer than 500 characters, ensuring conciseness for fast training, token-efficient fine-tuning, and mobile-friendly model development. The data format is JSON objects, each line containing an input field (Pashto prompt) and an output field (Pashto-translated psychology answer), e.g., input د اضطراب نښې څه دي? (What are the symptoms of anxiety?) and output the corresponding psychological explanation. The dataset size ranges from 1,000 to 10,000 samples, falling under text generation and translation tasks, licensed under Apache 2.0 for research and commercial use. Intended applications include building Pashto instruction-tuning models, mental health chatbots, domain-specific reasoning systems, and supporting fine-tuning for models like Qwen, LLaMA, Ministral, and Gemma. Note that this dataset is not a substitute for professional psychological advice, translations may have subtle stylistic variations, and it is recommended to combine with safety-aligned datasets during deployment to enhance model security.




