SCRum-9: Multilingual Stance Classification over Rumours on Social Media
收藏资源简介:
We release SCRum-9, the current largest multilingual rumour stance classification dataset as described in our paper "SCRum-9: Multilingual Stance Classification over Rumours on Social Media" (to appear in ICWSM 2026). The SCRum-9 dataset contains 7516 source-reply tweet pairs across 9 languages: English, Russian, Polish, French, Portuguese, Czech, Hindi, German and Spanish. The stance of the reply tweet towards a rumour-related source tweet is categorised into four stances: support, deny, query and comment. Each annotator indicates a first-choice stance label along with a confidence rating on a 5-point Likert scale (1 = extremely uncertain, 5 = absolutely certain). If the confidence rating is three or below, annotators are required to provide a second-choice stance label. We release the raw annotations from all the annotators (separated by "|" in the dataset file). If you use our dataset in your research, please cite: @article{li2025scrum, title={SCRum-9: Multilingual Stance Classification over Rumours on Social Media}, author={Li, Yue and Vasilakes, Jake and Zhao, Zhixue and Scarton, Carolina}, journal={arXiv preprint arXiv:2505.18916}, year={2025} }



