llm-japanese-dataset v0
收藏资源简介:
本研究构建了名为llm-japanese-dataset v0的日语聊天数据集,由东京大学开发,旨在为大型语言模型(LLMs)提供日语训练数据。该数据集包含约840万条记录,涵盖翻译、知识等多种任务类型。创建过程中,研究者整合了多个现有数据集,并通过特定格式进行处理。此数据集主要应用于提升非英语语言,特别是日语在LLMs中的处理能力,解决现有模型在非英语语言处理上的不足。
This study developed a Japanese chat dataset named llm-japanese-dataset v0, which was constructed by The University of Tokyo. The dataset is designed to provide Japanese training data for Large Language Models (LLMs). It contains approximately 8.4 million records, covering multiple task types including translation and knowledge-related tasks. During its development, researchers integrated multiple existing datasets and processed them in a specific format. This dataset is primarily used to enhance the processing capabilities of non-English languages, especially Japanese, in LLMs, addressing the shortcomings of current models in handling non-English languages.



