ytu-ce-cosmos/guardrail-tr
收藏资源简介:
Guardrail-TR是一个大规模土耳其语提示安全数据集,专门用于训练和评估内容审核或防护分类器。数据集包含约405,000个单轮用户提示,每个提示具有二元标签(安全或不安全)和多标签危害类别(一个提示可能属于多个类别)。所有发布的提示文本均为土耳其语;源自英语的行经过处理、翻译和文化适应性调整后纳入。数据集结构包括以下列:prompt(输入文本,土耳其语)、safety(安全或不安全)、category(危害类别列表,包含10个类别,如暴力犯罪、非暴力犯罪、仇恨歧视等,多标签允许)、language(行源语言,土耳其语或英语,但提示文本均为土耳其语)和source(上游数据集来源)。数据集基于MLCommons危害分类法进行标注,涵盖10个具体危害类型。数据来源于多个土耳其语和英语数据集,经过去重、启发式质量过滤、翻译、文化适应等处理流程,并通过LLM作为法官进行多标签标注和合成增强。数据集分为训练、验证和测试集,比例约为80%/10%/10%。其预期用途包括训练和评估土耳其提示安全分类器,研究多标签危害分类、过度拒绝(困难负样本)、注入/越狱检测等相关内容审核主题。同时,数据集存在有害内容、合成和翻译文本、分类法依赖、类不平衡和来源偏见等限制。
Guardrail-TR is a large-scale Turkish prompt-safety dataset for training and evaluating content-moderation or guardrail classifiers. The dataset contains approximately 405,000 single-turn user prompts with a binary safe/unsafe label and multi-label hazard categories (a prompt may carry more than one category). All released prompt text is Turkish; rows that originated from English sources were processed, translated, and culturally adapted before inclusion. The dataset structure includes the following columns: prompt (input text in Turkish), safety (safe or unsafe), category (a list of hazard labels from 10 categories, such as violent crimes, non-violent crimes, hate discrimination, etc., with multi-label allowed), language (source-language origin of the row, Turkish or English, but the prompt text is Turkish in either case), and source (upstream dataset origin). The labeling is based on the MLCommons hazards taxonomy, covering 10 specific hazard types. Data is sourced from multiple Turkish and English datasets, processed through deduplication, heuristic quality filtering, translation, cultural adaptation, and labeled using an LLM-as-judge approach with synthetic augmentation for underrepresented classes. The dataset is split into train, validation, and test sets with an approximate ratio of 80%/10%/10%. Its intended use includes training and evaluating Turkish prompt safety classifiers and research on multi-label hazard taxonomies, over-refusal (hard negatives), injection/jailbreak detection, and related moderation topics. Limitations include harmful content, synthetic and translated text, taxonomy and policy dependence, class imbalance, and source bias.




