CantHarm: A Cantonese Form-Sense-Audio Harmfulness Benchmark
收藏资源简介:
CantHarm is a Cantonese spoken safety benchmark release with 4,823 forms, 6,365 linked senses, 4,823 canonical audio clips, and 11 speaker IDs. The release includes harm labels, severity scores, public documentation, checksums, and audio packages. The dataset contains harmful/offensive lexical material and waveform audio and is intended for research and benchmark analysis. Audio files are licensed under CC BY 4.0 with speaker consent. Acceptable-use and ethics guidance discourages speaker recognition, voice cloning, biometric modeling, re-identification, and demographic inference; this guidance is separate from the CC BY 4.0 license grant. If older file-level Markdown documents in this record contain staging/pending wording, earlier wording about audio licensing, benchmark-scope wording, or non-release materials, this public record metadata and the current reviewer-facing documentation on the official website, GitHub, HuggingFace, and OSF supersede those older file-level notes for the 2026-04-02 public release. DOI: 10.5281/zenodo.20511573.



