OpenVoiceOS/yes_no_answers
收藏资源简介:
Yes/No多语言回答数据集是一个包含10,709个对话话语的多语言数据集,专门用于对是/否/模糊回答进行分类,覆盖43种语言。每个样本都是一个人可能对是/否问题做出的自然语言回答,分为三个类别:yes(肯定、同意或确认)、no(否定、拒绝或不同意)和None(真正模糊,没有上下文无法解决)。数据集还包括28个语义子类型,如直接肯定、强调否定、纯不确定性等,每个子类型在每个语言中至少有8个样本。所有话语都是通过大型语言模型(Claude)作为多语言对话AI直接生成,没有使用机器翻译,确保每个话语在目标语言中都是地道的。数据集支持多种语言,包括欧洲语言(如英语、德语、法语、西班牙语等)和亚洲及中东语言(如日语、韩语、中文、阿拉伯语等),并涵盖正式、中性和非正式语体。数据集经过全局去重处理,确保无重复条目,所有话语长度不超过75个字符,并保证文化真实性。
The Yes/No Multilingual Response Dataset is a multilingual dataset containing 10,709 conversational utterances, specifically designed for classifying yes/no/ambiguous responses across 43 languages. Each sample represents a natural language response that a person might give to a yes/no question, categorized into three classes: yes (affirmation, agreement, or confirmation), no (negation, refusal, or disagreement), and None (truly ambiguous and unsolvable without contextual information). The dataset also includes 28 semantic subtypes such as direct affirmation, emphatic negation, pure uncertainty, and more, with at least 8 samples per subtype for each language. All utterances were directly generated by Claude, a Large Language Model (LLM) acting as a multilingual conversational AI, without using machine translation, ensuring that each utterance is idiomatic in its target language. The dataset supports a wide range of languages, including European languages (e.g., English, German, French, Spanish, etc.) as well as Asian and Middle Eastern languages (e.g., Japanese, Korean, Chinese, Arabic, etc.), and covers formal, neutral, and informal speech registers. The dataset has undergone global deduplication to ensure no duplicate entries, all utterances are no longer than 75 characters, and cultural authenticity is guaranteed.



