OpceanAI/sota-hard
收藏资源简介:
SOTA Hard Reasoning Dataset是一个多语言、指令调优的推理数据集,支持英语、西班牙语、法语、中文、意大利语、俄语、日语、阿拉伯语、印地语、德语和捷克语。它包含机器生成和专家生成的内容,规模在10万到100万样本之间。该数据集设计用于多种任务,包括文本生成、翻译、问答、强化学习和文本分类,特别强调推理能力和思维链(Chain-of-Thought)方法,以促进复杂问题的解决。
The SOTA Hard Reasoning Dataset is a multilingual, instruction-tuned reasoning dataset that supports English, Spanish, French, Chinese, Italian, Russian, Japanese, Arabic, Hindi, German, and Czech. It includes both machine-generated and expert-generated content, with a size ranging from 100,000 to 1,000,000 samples. Designed for various tasks such as text generation, translation, question-answering, reinforcement learning, and text classification, the dataset emphasizes reasoning capabilities and Chain-of-Thought methods to facilitate complex problem-solving.



