遇见数据集

MTD-EU-26: Multilingual English–Urdu Text Detoxification Dataset

收藏
Zenodo2026-08-18 更新2026-08-20 收录
官方服务:

资源简介:

MTD-EU-26 is a multilingual text detoxification dataset containing 12,045 toxic social-media posts/comments and their corresponding detoxified rewrites in English and Urdu. The corpus consists of 9,019 English instances collected from Reddit, X (formerly Twitter), and YouTube, and 3,026 Urdu instances collected from Facebook and X. Each toxic instance is paired with a detoxified rewrite in the same language, with the aim of removing toxic or offensive expressions while preserving the original meaning. The English and Urdu subsets were collected independently and are not translations of each other. The dataset was developed to support research on multilingual and low-resource text detoxification, toxicity mitigation, meaning preservation, and safe natural language processing.

提供机构:
Zenodo
创建时间:
2026-08-18
二维码
社区交流群
二维码
科研交流群
商业服务