政治播客毒性数据集
收藏资源简介:
该数据集由印度理工学院克勒格布尔分校的研究团队创建,旨在分析美国政治播客中的毒性内容。数据集包含31个流行政治播客的转录和对话分析,涵盖了2022-2023年间的音频数据。通过Whisper模型进行转录,并使用Pyannote进行说话人分离,最终生成了8634条毒性对话链。数据集的应用领域包括毒性内容检测、对话结构分析以及实时监控和干预机制的开发,旨在解决播客中的毒性传播问题,促进健康的公共讨论。
This dataset was developed by a research team from the Indian Institute of Technology Kharagpur, with the goal of analyzing toxic content in U.S. political podcasts. It comprises transcriptions and conversational analyses of 31 popular political podcasts, covering audio data collected between 2022 and 2023. Transcriptions were generated using the Whisper model, and speaker diarization was conducted via Pyannote, ultimately producing 8,634 toxic conversation chains. Application domains of this dataset include toxic content detection, conversational structure analysis, and the development of real-time monitoring and intervention mechanisms. The dataset aims to mitigate toxic speech propagation in podcasts and promote healthy public discourse.

- 1Dynamics of Toxicity in Political Podcasts印度理工学院克勒格布尔分校 · 2025年



