Ultra Chat 200k Dutch
收藏资源简介:
Ultra Chat 200k Dutch数据集是由鲁汶大学和荷兰语言研究所创建的高质量对话数据集,旨在提升荷兰语生成语言模型的对话能力。该数据集包含192,598条对话,涵盖了技术、艺术、创业等多个主题,通过GPT-4生成,强调了多样性和覆盖范围。数据集的创建过程包括使用GPT-4进行多轮对话生成,并模拟不同用户角色以增加数据的多样性。该数据集主要应用于荷兰语生成语言模型的微调,旨在解决荷兰语对话模型质量不足的问题,提升荷兰语用户的技术体验。
The Ultra Chat 200k Dutch Dataset is a high-quality conversational dataset developed by KU Leuven and the Dutch Language Institute, designed to enhance the conversational capabilities of Dutch generative language models. It contains 192,598 dialogues spanning a wide range of topics including technology, art, entrepreneurship and other fields, and was generated using GPT-4, with significant emphasis placed on data diversity and coverage. The dataset creation process includes generating multi-turn conversations via GPT-4, as well as simulating different user roles to further boost data diversity. This dataset is primarily applied to the fine-tuning of Dutch generative language models, with the goal of addressing the issue of insufficient quality of current Dutch conversational models and improving the technical experience for Dutch users.

- 1GEITje 7B Ultra: A Conversational Model for Dutch鲁汶大学,荷兰语言研究所 · 2024年



