The Punjabi Geet Toxicity Dataset (PGT-FG): Fine-Grained Multimodal Annotations for Toxic Content in Punjabi Songs
收藏资源简介:
PGT-FG is a fine-grained multimodal dataset to analyze toxic content in Punjabi songs. The dataset consists of 2,396 unique 30-second segments from 397 Punjabi songs. Each segment is annotated with binary toxicity, fine-grained toxicity categories, severity, sentiment, mode, and annotation uncertainty. To evaluate inter-annotator agreement, a common set of 158 segments was independently annotated by 4 annotators. The public release includes resolved segment-level annotations, song and segment metadata, source YouTube references, documentation, and supporting code where applicable. The annotation process resulted in 2,870 annotation records. The public release does not include those original annotator-level records. In the released annotation file, each of the 2,396 unique segments has exactly one resolved record. No copyrighted audio, video, or full lyric text is redistributed. For reference to the original source material, YouTube URLs and time stamps are provided.



