quannguyen204/diasynth-vi-elderly-orpo-v2
收藏资源简介:
DiaSynth越南老年人护理—ORPO数据集是一个用于ORPO微调的偏好对数据集,专注于越南老年人护理安全风险(如电话诈骗、用药错误等)。每条记录包含单轮对话的(prompt, chosen, rejected)三元组,其中chosen是安全且符合角色设定(助手使用con自称)的响应,rejected是对同一风险提示的不安全或偏离角色设定的变体。数据集包含训练集(14,927对)、验证集(1,494对)和测试集(165对),并经过Unicode标准化、长度过滤、去重和分层划分等预处理。该数据集旨在通过ORPO微调(使用β=0.1)在老年人护理风险下对齐模型的安全性和角色一致性,但数据为合成生成,非人工标注标准,且覆盖范围仅限于老年人领域。
DiaSynth Vietnamese Elderly Care — ORPO is a preference pairs dataset for ORPO finetuning, focused on Vietnamese elderly-care safety hazards (e.g., phone scams, medication mistakes). Each record is a single-turn (prompt, chosen, rejected) tuple where chosen is a safe and persona-consistent response (assistant uses con for self-reference) and rejected is an unsafe or off-persona variant for the same hazard prompt. The dataset includes train (14,927 pairs), validation (1,494 pairs), and test (165 pairs) splits, with preprocessing such as NFC Unicode normalization, token length filtering, deduplication, and stratified splitting. It is intended for ORPO finetuning (with β=0.1) to align model safety and persona under elderly-care hazards, but data is synthetic (not human-labeled gold standard) and coverage is biased toward the elderly domain.




