遇见数据集

CantHarm: A Cantonese Form-Sense-Audio Harmfulness Benchmark

收藏
Zenodo2026-06-04 更新2026-06-05 收录
官方服务:

资源简介:

CantHarm is a Cantonese spoken safety benchmark release with 4,823 forms, 6,365 linked senses, 4,823 canonical audio clips, and 11 speaker IDs. The release includes harm labels, severity scores, public documentation, checksums, and audio packages. The dataset contains harmful/offensive lexical material and waveform audio and is intended for research and benchmark analysis. Audio files are licensed under CC BY 4.0 with speaker consent. Acceptable-use and ethics guidance discourages speaker recognition, voice cloning, biometric modeling, re-identification, and demographic inference; this guidance is separate from the CC BY 4.0 license grant. If older file-level Markdown documents in this record contain staging/pending wording, earlier wording about audio licensing, benchmark-scope wording, or non-release materials, this public record metadata and the current reviewer-facing documentation on the official website, GitHub, HuggingFace, and OSF supersede those older file-level notes for the 2026-04-02 public release. DOI: 10.5281/zenodo.20511573.

提供机构:
Zenodo
创建时间:
2026-06-02
二维码
社区交流群
二维码
科研交流群
商业服务