CantHarm: A Cantonese Form-Sense-Audio Harmfulness Benchmark
收藏资源简介:
CantHarm is a Cantonese spoken safety benchmark release with 4,823 forms, 6,365 linked senses, 4,823 canonical audio clips, and 11 speaker IDs. The release includes harm labels, severity scores, public documentation, checksums, and audio packages. The dataset contains harmful/offensive lexical material and waveform audio and is intended for research and benchmark analysis. Audio files are licensed under CC BY 4.0 with speaker consent. Acceptable-use and ethics guidance discourages speaker recognition, voice cloning, biometric modeling, re-identification, and demographic inference; this guidance is separate from the CC BY 4.0 license grant. If older file-level Markdown documents in this record contain staging/pending wording, earlier wording about audio licensing, benchmark-scope wording, or non-release materials, this public record metadata and the current reviewer-facing documentation on the official website, GitHub, HuggingFace, and OSF supersede those older file-level notes for the 2026-04-02 public release. DOI: 10.5281/zenodo.20511573.
CantHarm是一款粤语口语安全基准发布数据集,包含4823个词形、6365个关联义项、4823条标准音频片段以及11位说话者身份标识。本次发布附带危害标签、严重程度评分、公开文档、校验和与音频数据包。本数据集包含有害/冒犯性词汇材料与波形音频,仅用于研究与基准分析。音频文件已获得说话者知情同意,采用CC BY 4.0许可协议进行授权。可接受使用规范与伦理准则明确禁止开展说话者识别、语音克隆、生物特征建模、重识别及人口统计推断工作;该准则独立于CC BY 4.0许可授权范围。若本记录中的旧版文件级Markdown文档包含暂存/待定稿表述、早期音频许可相关表述、基准范围相关表述或非发布材料,则针对2026年4月2日的公开发布版本,本公开记录元数据以及官方网站、GitHub、HuggingFace与OSF平台上面向审核人员的当前文档,将取代上述旧版文件级注释。数字对象标识符(DOI):10.5281/zenodo.20511573。



