遇见数据集

Multilingual Reddit Dataset on Global Discourse about Volodymyr Zelenskyy and Ukraine

收藏
DataONE2026-05-15 更新2026-05-27 收录
官方服务:

资源简介:

Overview This repository contains an optimally safe metadata-only multilingual dataset collected and processed as part of a study on public online discourse about Volodymyr Zelenskyy and Ukraine in Reddit communities. The dataset includes depersonalized and aggregated post-level metadata, query-language indicators, community-level information, engagement intervals, and sentiment labels. It is intended to support academic research on multilingual digital discourse, public opinion dynamics, sentiment distribution, information behavior, and online communication patterns related to Ukraine and Volodymyr Zelenskyy. The dataset was prepared by Yuriy Syerov and Solomiia Fedushko. The data were collected from public Reddit posts using multiple spelling and language variants of the name Volodymyr Zelenskyy, including Latin, Cyrillic, Japanese, and Chinese forms. The released version follows a strict data-minimization approach: it does not include Reddit usernames, user profile links, direct post URLs, full post texts, user-level karma indicators, membership-duration fields, exact timestamps, or user-profile community identifiers. Dataset reddit_zelenskyy_discourse_metadata_OPTIMAL_SAFE This dataset contains the optimally safe open metadata version of Reddit posts mentioning Volodymyr Zelenskyy across multiple query variants. The dataset was prepared for public repository access by removing direct and indirect user-identifying fields and by generalizing potentially sensitive metadata. Key Features: Record ID: Unique internal identifier for each dataset record. Query Term: The spelling or language variant used to identify relevant Reddit posts. Query Language: Language or script category associated with the query variant. Post Type: General type of Reddit content included in the dataset. Community Category: Depersonalized subreddit-level information, with user-profile communities removed or generalized. Community Size Interval: Approximate interval of community size, where available, instead of exact member counts. Engagement Intervals: Generalized intervals for upvotes, comments, and awards rather than exact values. Temporal Granularity: Month-level temporal information instead of exact date and time. Sentiment Label: General tone of the post, such as positive, neutral, or negative, inherited from the source export. Purpose: This metadata-only dataset is designed to support academic analysis of Reddit-based public discourse while minimizing privacy and re-identification risks. It can be used for aggregate-level analysis of engagement patterns, multilingual query coverage, community-level discourse distribution, sentiment dynamics, and temporal trends in discussions related to Ukraine and Volodymyr Zelenskyy. Research Context The dataset was developed to support research on how Volodymyr Zelenskyy and Ukraine are discussed in international online communities. Reddit was selected as a relevant platform because it contains diverse public discussions, multilingual user-generated interactions, and community-based communication structures. The dataset enables researchers to examine aggregate-level patterns of public discourse, sentiment, engagement, and community distribution during periods of geopolitical crisis and war-related communication. The research context includes digital public opinion, multilingual information flows, online community reactions, sentiment dynamics, and discourse patterns related to Ukraine and its political leadership. The dataset may be useful for studies in social communication, information science, computational social science, social media analytics, sentiment analysis, online community research, and multilingual digital discourse analysis. Data Minimization and Privacy Protection The released dataset follows an optimally safe metadata-only model. The following categories of information are not included in the public version: Reddit usernames, author profile URLs, direct post URLs, full post texts, user-level karma fields, membership-duration fields, exact timestamps, direct user-profile community names, and raw identifiers that could support tracing individual Reddit users. Potentially identifying or highly granular fields were removed, generalized, or converted into broader analytical intervals. Exact engagement values were replaced with intervals; exact date and time were reduced to month-level temporal information; and user-profile community references were removed or marked as removed. This approach supports aggregate academic analysis while reducing the risk of user identification, profiling, or inappropriate reuse. Ethical Considerations The dataset was prepared with attention to responsible data management and ethical reuse of social media-derived data. Although the source data were collected from publicly available Reddit posts, the public release was designed to avoid exposing individual Reddit users or enabling user-level tracing. The dataset is intended for academic research...

创建时间:
2026-05-18
二维码
社区交流群
二维码
科研交流群
商业服务