Persian_QA
收藏资源简介:
该数据集包含5900对波斯语问答,由AI助手生成,利用GPT-4o模型通过Avala API服务。答案具有标准波斯语掌握、科学准确、清晰易懂、实用例子、结构化、正确标点、诚实承认不确定性、提供可靠来源、根据受众调整、提供长答案摘要等特点。数据集适合训练波斯语模型、开发问答系统、进行自然语言处理任务和教育目的。数据格式为CSV,包含问题和答案两列。
This dataset contains 5,900 Persian question-answer pairs, generated by AI assistants using the GPT-4o model via the Avala API service. The generated answers boast standard Persian language proficiency, scientific accuracy, clarity and comprehensibility, practical examples, structured formatting, correct punctuation, honest acknowledgment of uncertainty, provision of reliable sources, adaptation to the target audience, and inclusion of summaries for lengthy answers. This dataset is suitable for training Persian language models, developing question-answering systems, conducting natural language processing tasks, and educational purposes. The data is stored in CSV format, with two columns: "question" and "answer".
Persian Question-Answer Dataset
数据集描述
该数据集包含5900对波斯语(Farsi)问答对,由PersianAnswerGenerator类从answer.py生成。答案由AI助手通过Avala API服务利用GPT-4o模型生成。
答案特点
- 标准波斯语的完全掌握
- 对问题的准确和科学回答
- 清晰易懂的解释
- 使用实际例子以更好地理解概念
- 结构化的回答,适当的段落划分
- 正确使用波斯语书写标点符号
- 在需要时诚实地承认不确定性
- 提供可靠的进一步学习资源
- 根据受众调整回答水平
- 为长答案提供总结
数据集生成过程
- 数据收集:加载了5900个从lmsys问题翻译过来的波斯语问题。
- 答案生成:通过Avala API与GPT-4o模型交互,生成详细且准确的答案。
- 数据处理:清理和组织生成的答案。
使用场景
该数据集适用于:
- 训练波斯语语言模型
- 开发问答系统
- 执行自然语言处理任务
- 教育目的
数据格式
数据集以CSV格式提供,包含两列:
question:波斯语问题answer:波斯语详细答案
许可证
该数据集在MIT许可证下发布。




