dpo-icelandic-interpreted
收藏资源简介:
该数据集是冰岛语版本的直接偏好优化(DPO)配对数据,源自原始数据集 argilla/ultrafeedback-binarized-preferences-cleaned。所有数据经过翻译和文化适应,以贴合冰岛语习惯和本地文化。生成过程中使用了 kimi-k3 模型,并采用专门的提示策略,避免产生直译腔或机械翻译。具体而言,该模型被指示将提示和首选回答高度本地化,同时刻意生成一个质量低下的拒绝回答,其中包含糟糕的翻译、过多的英语借词(冰英混杂)或美国中心幻觉。这种设计旨在为 DPO/RLHF 训练提供高质量、针对性强的信号,引导模型偏好地道、文化契合的冰岛语,而非死板的词典翻译。数据集包含以下字段:prompt(文化适应的冰岛语提示)、chosen(高质量地道冰岛语回答)、rejected(低质量翻译回答)、language(目标语言,此处为冰岛语)。许可证沿用源数据集的 MIT 许可证。
This dataset is an Icelandic version of Direct Preference Optimization (DPO) paired data, derived from the original dataset argilla/ultrafeedback-binarized-preferences-cleaned. All data has been translated and culturally adapted to align with Icelandic language habits and local culture. The generation process used the kimi-k3 model with a specialized prompting strategy to avoid translationese or machine translation. Specifically, the model was instructed to highly localize the prompt and chosen response, while deliberately generating a low-quality rejected response containing poor translation, excessive English loanwords (Icelandic-English mixing), or American-centric hallucinations. This design aims to provide high-quality, targeted signals for DPO/RLHF training, guiding the model to prefer natural, culturally appropriate Icelandic over rigid dictionary translations. The dataset includes the following fields: prompt (culturally adapted Icelandic prompt), chosen (high-quality, natural Icelandic response), rejected (low-quality translated response), language (target language, here Icelandic). The license follows the original datasets MIT license.
DPO Icelandic Interpreted 数据集概述
数据集基本信息
- 数据集地址:https://huggingface.co/datasets/jonasaise/dpo-icelandic-interpreted
- 语言:冰岛语(is)
- 许可证:MIT
- 标签:dpo、rlhf、ultrafeedback、synthetic、kimi-k3、icelandic
- 配置名称:default
- 数据文件路径:data/**/*.parquet
- 数据划分:train
数据集描述
该数据集包含从原始数据集 argilla/ultrafeedback-binarized-preferences-cleaned 翻译并进行文化适配为冰岛语的直接偏好优化(DPO)配对数据。
来源与方法论
- 源数据集:argilla/ultrafeedback-binarized-preferences-cleaned
- 用于合成的模型:kimi-k3(通过本地端点)
- 翻译策略:数据集使用专门设计的提示词生成,以防止“翻译腔”和直接直译。kimi-k3 模型被指示对提示词和被选回答进行深度本地化,以贴合冰岛文化和习语。此外,该模型特意生成一个被拒回答,该回答存在翻译质量差、过多英语外来词(Swenglish/Ice-lish)或美国中心主义的幻觉内容。这为 DPO/RLHF 提供了高质量、有针对性的信号,用于教导模型偏好真实、具有文化根基的冰岛语,而非机械的字典式翻译。
数据结构
数据集包含以下字段:
prompt:经过文化适配的冰岛语提示词。chosen:高质量、地道的冰岛语回答。rejected:翻译质量差、未经校准的回答。language:该行的目标语言。
许可证
该数据集继承其源数据集 argilla/ultrafeedback-binarized-preferences-cleaned 的许可证(MIT)。





