DetoxLLM
收藏资源简介:
DetoxLLM数据集是一个用于文本生成任务中去毒化的数据集。它包含多个特征,如data_id、toxic、non_toxic、explanation、platform和source_label。数据集分为训练集、验证集和测试集,分别包含7453、2041和955个样本。数据集的创建使用了ChatGPT生成跨平台的伪并行去毒化数据。其目的是帮助研究人员构建一个端到端的去毒化框架,并作为一个有前景的基线来开发更强大和有效的去毒化框架。然而,数据集也存在一些局限性,如数据生成过程依赖于ChatGPT,数据质量可能存在问题,模型响应可能不完全保留原意,以及潜在的伦理风险和偏见。
The DetoxLLM dataset is a benchmark dataset for detoxification in text generation tasks. It includes multiple features such as data_id, toxic, non_toxic, explanation, platform, and source_label. The dataset is split into training, validation, and test sets, containing 7453, 2041, and 955 samples respectively. The dataset was constructed using pseudo-parallel detoxification data generated by ChatGPT across multiple platforms. Its core objective is to help researchers build end-to-end detoxification frameworks, and act as a promising baseline for developing more robust and effective detoxification systems. However, the dataset also has several limitations: its data generation process relies on ChatGPT, which may lead to potential data quality issues, model responses may not fully preserve the original intent, and there exist potential ethical risks and biases.




