conversationalquestion-answer-wikipedia-v1.0
收藏资源简介:
该数据集名为'conversationalquestion-answer-wikipedia-v1.0',由Restack机构创建,包含10000个问答对,数据来源于Wikipedia文章。该数据集用于训练语言模型,使其在语音交互中能够以自然、对话式的语气进行回答。数据集的创建过程包括从Wikipedia文章中提取文本段落,并使用语言模型生成问题和答案,然后通过Flesch阅读流畅度评分筛选出符合对话式语气的问题和答案。该数据集适用于解决语音交互中语言模型回答风格的问题,为开发语音助手等应用提供支持。
This dataset, named 'conversationalquestion-answer-wikipedia-v1.0', was created by the organization Restack. It comprises 10,000 question-answer pairs sourced from Wikipedia articles. This dataset is designed for training language models to generate natural, conversational responses during voice interactions. The dataset creation workflow includes extracting text passages from Wikipedia articles, generating corresponding question-answer pairs via language models, and filtering the pairs to meet conversational tone standards using the Flesch Reading Ease score. It addresses the issue of non-standardized response styles of language models in voice interactions, and provides support for developing applications such as voice assistants.




