arsyra-chatbot
收藏资源简介:
ArSyra Chatbot数据集是一个专为阿拉伯语对话AI系统设计的训练数据集,旨在优化聊天机器人、虚拟助手和对话系统的性能。数据集包含1,297个经过质量筛选的阿拉伯语对话样本,涵盖自然对话对、问候与告别模式、指令遵循示例、自由形式开放回答以及正式与非正式语体转换等多种对话类型。所有数据均来自母语为阿拉伯语的用户与结构化提示的互动,反映了真实的阿拉伯语交流模式,并覆盖多种方言群体。数据集包含多个字段,如文本内容、类别、国家、方言群体、质量评分等,适用于文本生成和对话AI等任务。数据以CC-BY-NC-SA-4.0许可证发布,提供50个样本的预览版本,完整数据集需申请获取。
The ArSyra Chatbot Dataset is a training dataset specifically designed for Arabic conversational AI systems, aimed at optimizing the performance of chatbots, virtual assistants, and dialogue systems. The dataset contains 1,297 quality-filtered Arabic conversational samples, covering a wide range of dialogue types including natural dialogue pairs, greeting and farewell patterns, instruction-following examples, free-form open-ended responses, and formal-informal stylistic shifts. All data originate from interactions between native Arabic speakers and structured prompts, reflecting authentic Arabic communication patterns and spanning multiple Arabic dialect groups. The dataset comprises multiple fields such as text content, category, country of origin, dialect group, quality score, among others, and is suitable for tasks including text generation and conversational AI. The dataset is distributed under the CC-BY-NC-SA-4.0 license. A preview version with 50 samples is provided, and access to the full dataset requires formal application.



