Calling Out Antisemitism: A Dataset for Distinguishing Between Antisemitic and Counter-Antisemitic Discourse
收藏资源简介:
Calling Out Antisemitism: A Dataset for Distinguishing Between Antisemitic and Counter-Antisemitic Discourse A Dataset by the Social Media & Hate Research Lab, Institute for the Study of Contemporary Antisemitism (ISCA), Borns Jewish Studies Program, Indiana University Bloomington Dataset Overview This dataset represents a re-annotation of a subset of tweets from the dataset Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics. The original dataset (Jikeli et al., 2023) contained 11,311 tweets collected between January 2019 and April 2023, labeled for antisemitic content following the IHRA working definition of antisemitism. The current release focuses on 894 tweets from two time periods that were re-annotated to identify “calling out antisemitism”—tweets that explicitly condemn, counter, or discuss antisemitic language or expressions. Jews (September–December 2022): 459 tweets, including 140 calling out antisemitism (30.50%) and 88 antisemitic tweets (19.17%). Israel (January–April 2023): 435 tweets, including 42 calling out antisemitism (9.66%) and 87 antisemitic tweets (20.00%). Combined, these subsets provide additional annotation layers that capture “calling out antisemitism” — tweets that raise awareness about, condemn, or counter antisemitic content within antisemitism-related conversations on X (formerly Twitter). Relation to the Original Dataset The original dataset included random samples of tweets containing relevant keywords (“Jews,” “Israel,” “ZioNazi*,” “K---s”) between January 2019 and April 2023. It featured annotation for antisemitism, bias, and related discourse dimensions. This new release extends that work by adding a “Calling Out” label to better capture the ways users engage with or challenge antisemitic content. Annotation Process Annotation was performed via the AnnotHate Portal, enabling annotation in full conversational and visual context (including images, quoted material, and replies). All tweets were annotated by two trained annotators from the Social Media & Hate Research Lab. Discrepancies were discussed collaboratively to ensure high inter-annotator reliability. The “Calling Out” category was used to distinguish tweets that actively call out or condemn antisemitism from those that propagate it. These subsets were re-annotated by a trained expert to ensure consistency and reliability. File Description The dataset is provided in CSV format, where each row represents a single tweet. Columns include: id: Unique identifier for each tweet created_at: Timestamp of publication Biased: Binary label indicating whether the tweet is antisemitic (1) or not (0) Keyword: Search keyword used in the query (“Jews” or “Israel”) text: Full, unprocessed tweet text CallingOut: Binary label indicating whether the tweet calls out or condemns antisemitism (1) or not (0) Acknowledgements This work utilized Jetstream2 at Indiana University through allocation HUM200003 from the Advanced Cyberinfrastructure Coordination Ecosystem: Services & Support (ACCESS) program, supported by the U.S. National Science Foundation (grants #2138259, #2138286, #2138307, #2137603, and #2138296). References Jikeli, G., Karali, S. Miehling, D. Soemer, K. (2023): Antisemitic Messages? A Guide to High-Quality Annotation and a Labeled Dataset of Tweets. https://arxiv.org/abs/2304.14599 Jikeli, G., Karali, S., Miehling, D., & Soemer, K. (2024). Antisemitism on Twitter: A Dataset for Machine Learning and Text Analytics [Data set]. Zenodo. https://doi.org/10.5281/zenodo.14448399



