遇见数据集

Twitter Dataset on Russo-Ukrainian War with Multi-Model Annotations

收藏
Zenodo2026-02-18 更新2026-05-26 收录
官方服务:

资源简介:

Description This dataset was created as part of an independent exploratory research project focused on information warfare and NLP applications in conflict analysis. This dataset consists of obfuscated Twitter posts related to the Russo-Ukrainian war, collected over the period of the years 2022-2023. The data has been processed to remove personally identifiable information (PII) and sensitive content while retaining semantic and contextual integrity. Each post is annotated with predictions from five state-of-the-art NLP models for various tasks, including: Sentiment Analysis Propaganda Detection Emotion Classification etc. These models offer complementary perspectives on the content of the posts and are useful for research in computational social science, conflict analysis, misinformation studies, and more. Dataset Structure The dataset is provided in CSV format and contains the following fields: Field Name Description date_month Month of post publication (format: YYYY-MM), used for temporal analysis. replyCount Number of replies to the tweet. retweet_count Number of times the tweet was retweeted. like_count Number of likes (favorites) the tweet received. quote_count Number of quote tweets referencing the post. lang Detected language of the tweet (ISO 639-1 code, e.g., en, uk, ru). source_label Platform or app used to publish the tweet (e.g., "Twitter for Android"). hashtags Comma-separated list of hashtags used in the tweet. place_generalized Generalized or anonymized geolocation data (e.g., country or region level). urls_filtered Cleaned list of URLs included in the tweet, if any. sentiment Sentiment label predicted by the sentiment analysis model (positive = 1, neutral = 0, negative = -1). emotion Emotion label from the emotion classification model (anger, sadness, joy, fear, etc.). propaganda_binary Binary classification indicating whether propaganda was detected (1 = propaganda, 0 = not). propaganda_18 Specific propaganda technique label based on a 18-class taxonomy. fake_news Binary label indicating likelihood of the tweet spreading fake or misleading information (1 = likely fake, 0 = likely factual). Obfuscation Policy All usernames, hashtags, and URLs have been replaced or removed No direct quote or identifier linking to individuals remains

提供机构:
Zenodo
创建时间:
2025-04-18
二维码
社区交流群
二维码
科研交流群
商业服务