MTD-EU-26: Multilingual English–Urdu Text Detoxification Dataset
收藏资源简介:
MTD-EU-26 is a multilingual text detoxification dataset containing 12,045 toxic social-media posts/comments and their corresponding detoxified rewrites in English and Urdu. The corpus consists of 9,019 English instances collected from Reddit, X (formerly Twitter), and YouTube, and 3,026 Urdu instances collected from Facebook and X. Each toxic instance is paired with a detoxified rewrite in the same language, with the aim of removing toxic or offensive expressions while preserving the original meaning. The English and Urdu subsets were collected independently and are not translations of each other. The dataset was developed to support research on multilingual and low-resource text detoxification, toxicity mitigation, meaning preservation, and safe natural language processing.



