polite-guard
收藏资源简介:
Polite Guard数据集是一个用于文本分类任务的合成和注释数据集,主要任务是将文本分类为礼貌、有些礼貌、中立和不礼貌四个类别。数据集由50,000个通过Few-Shot提示生成的样本、50,000个通过Chain-of-Thought提示生成的样本以及200个来自企业培训的注释样本组成。合成数据被分为训练集(80%)、验证集(10%)和测试集(10%),每个集合都根据标签进行了平衡。注释数据仅用于评估。每个样本包含文本输入、分类标签、生成文本的语言模型来源以及生成文本时的推理过程。数据集涵盖了多个行业的客户服务互动,包括金融、旅游、餐饮、零售、体育俱乐部、文化和教育以及专业发展。
The Polite Guard Dataset is a synthetic and annotated dataset tailored for text classification tasks, with its core task being to categorize texts into four classes: polite, somewhat polite, neutral, and impolite. This dataset comprises 50,000 samples generated via Few-Shot prompting, 50,000 samples generated via Chain-of-Thought prompting, and 200 annotated samples sourced from enterprise training materials. The synthetic data is split into training (80%), validation (10%), and test (10%) sets, with each set balanced in accordance with its label distribution. The annotated samples are solely used for evaluation purposes. Each sample includes the text input, classification label, the source large language model (LLM) used to generate the text, and the reasoning process applied during text generation. The dataset covers customer service interactions across multiple industries, including finance, tourism, catering, retail, sports clubs, culture and education, and professional development.




