UPDESH1
收藏资源简介:
UPDESH1是一个高质量的、大规模的合成指令遵循数据集,包含950万条数据点,覆盖13种印度语言。它包括多样化的推理和生成任务,重点在于增强长上下文和多轮对话能力,并提高与印度文化环境的契合度。数据集的创建采用了一种自下而上的生成策略,通过提示大型开源LLMs(参数≥235B)将数据生成扎根于语言特定的维基百科内容。评估结果显示,生成的数据质量高,但在人类评估中也指出了需要改进的特定领域。
UPDESH1 is a high-quality, large-scale synthetic instruction-following dataset containing 9.5 million data points spanning 13 Indian languages. It encompasses diverse reasoning and generation tasks, with a focus on enhancing long-context and multi-turn conversation capabilities, as well as improving alignment with Indian cultural contexts. The dataset was developed using a bottom-up generation strategy, where large open-source LLMs (with parameters ≥235B) are prompted to generate data grounded in language-specific Wikipedia content. Evaluation results demonstrate that the generated data is of high quality, though human evaluators have also identified specific domains that require further improvement.




