Guardrail-TR是一个大规模土耳其语提示安全数据集,专为训练和评估内容审核/护栏分类器而构建。该数据集包含约405,000个单轮用户提示,所有提示文本均以土耳其语发布。每个数据样本包含一个二元安全标签(“safe”或“unsafe”)以及一个多标签危害类别列表,允许一个提示对应多个危害类别。危害分类体系包含10个类别,灵感来源于MLCommons危害分类法,具体包括:暴力犯罪、非暴力犯罪、仇恨歧视、骚扰冒犯、成人性内容、儿童性虐待/剥削、自残自杀、提示注入/越狱、虚假信息/政治操纵、隐私侵犯。数据集结构清晰,包含提示文本、安全标签、危害类别列表、源语言标识(‘tr’或‘en’)以及上游数据来源等字段。数据源自多个土耳其语和英语的开源数据集,并经过严格的去重、质量过滤、翻译(针对英语源)、文化适应、LLM陪审团标注(使用三个大模型进行多标签标注)以及针对少数类别的合成增强等处理流程。最终数据按约80%/10%/10%的比例划分为训练集、验证集和测试集。该数据集主要用于土耳其语提示安全分类模型的开发与评估,以及多标签危害分类、过度拒绝(困难负样本)、注入/越狱检测等相关内容审核研究。需要注意的是,数据集包含具有攻击性、暴力、性相关、自残等性质的有害文本,使用时需谨慎。同时,数据集中存在大量机器翻译和文化适应的文本,且标注依赖于模型判断,可能存在残留的翻译痕迹、文化错配或判断误差。数据集遵循Apache 2.0许可证。
Guardrail-TR is a large-scale Turkish prompt safety dataset built specifically for training and evaluating content moderation/guardrail classifiers. The dataset contains approximately 405,000 single-turn user prompts, all in Turkish. Each data sample includes a binary safety label ("safe" or "unsafe") and a multi-label harm category list, allowing one prompt to correspond to multiple harm categories. The harm classification taxonomy consists of 10 categories, inspired by the MLCommons harm taxonomy, specifically including: Violent Crime, Non-violent Crime, Hatred and Discrimination, Harassment and Offensive Content, Adult Content, Child Sexual Abuse/Exploitation, Self-harm and Suicide, Prompt Injection/Jailbreaking, Misinformation/Political Manipulation, and Privacy Violation. The dataset has a clear structure, containing fields such as prompt text, safety label, harm category list, source language identifier ('tr' or 'en'), and upstream data source. The data is sourced from multiple Turkish and English open-source datasets, and has undergone strict processing workflows including deduplication, quality filtering, translation (for English sources), cultural adaptation, LLM jury annotation (multi-label annotation using three large language models), and synthetic augmentation for minority classes. The final dataset is split into training, validation, and test sets at a ratio of approximately 80%/10%/10%. This dataset is primarily used for the development and evaluation of Turkish prompt safety classification models, as well as related content moderation research such as multi-label harm classification, over-rejection (hard negative samples), and injection/jailbreak detection. It should be noted that the dataset contains harmful text with characteristics of aggression, violence, sexual content, self-harm and other similar natures, and should be used with caution. In addition, the dataset contains a large amount of machine-translated and culturally adapted text, and the annotation relies on model judgments, which may have residual translation traces, cultural mismatches or judgment errors. The dataset is licensed under Apache 2.0.