Labeled Datasets for Research on Information Operations
收藏资源简介:
Labeled Datasets for Research on Information Operations Compliance with Platform Terms To comply with the platform terms, we ask that you download one data file per researcher, per day. README 19-November-2024Contact: Observatory on Social Media Dataset ArticlesThis dataset is collected and processed according to the paper "Labeled Datasets for Research on Information Operations." DescriptionThese datasets contain data curated for research on information operations (IO) and includes both labeled IO and control data. The datasets cover 26 verified IO campaigns from various countries and provide comprehensive records of posts from IO accounts alongside control posts from legitimate accounts discussing similar topics during the same periods. The datasets enable the development and benchmarking of IO detection methods by comparing coordinated versus organic accounts. LicenseThis dataset is available under the Attribution-NonCommercial-NoDerivatives 4.0 International license. If you use this data, please cite the original paper. Dataset ContentThe dataset includes anonymized fields to preserve privacy, and is structured with the following columns: postid: Unique identifier for each post within the dataset. post_text: The textual content of the post. The PII inside post_text such as mentions and URLs are hashed application_name: Hashed version of the name of the application or platform from which the post was made. post_language: Language in which the post was written. in_reply_to_postid: Anonymized ID of the post this entry is replying to, if applicable. in_reply_to_accountid: Anonymized ID of the account the post is replying to, if applicable. post_time: Timestamp indicating when the post was made. accountid: Unique anonymized ID for the account that created the post. account_profile_description: Description provided by the account holder in their profile. follower_count: Number of followers the account had at the time of data collection. following_count: Number of accounts the user was following at the time of data collection. account_creation_date: Date when the account was created. is_repost: Boolean indicator if the post is a repost. reposted_accountid: Anonymized ID of the original account that made the reposted post, if applicable. reposted_postid: Anonymized ID of the original post that was reposted, if applicable. hashtags: Hashtags included in the post content, if any. urls: Hashed URLs shared within the post, if any. account_mentions: Anonymized ID of accounts mentioned within the post, if any. is_control: Boolean indicator marking whether the post is from a control (True) or IO (False) account. Data for different campaigns are organized in separate versions of this repository.
用于信息作战(Information Operations, IO)研究的标注数据集 ## 平台条款合规要求 为遵守平台相关条款,我们要求每位研究者每日仅可下载一份数据文件。 ## 自述文件 2024年11月19日 | 联系方式:社交媒体观测站(Observatory on Social Media) ### 数据集关联论文 本数据集依据《用于信息作战研究的标注数据集》论文进行采集与处理。 ### 数据集说明 本数据集专为信息作战研究整理,包含标注的信息作战数据与对照数据。数据集涵盖来自多个国家的26个经核实的信息作战行动,完整记录了信息作战账号发布的帖文,以及同期合法账号发布的同类主题对照帖文。通过对比协同账号与自然账号的行为,本数据集可用于信息作战检测方法的开发与性能基准测试。 ### 授权许可 本数据集采用署名-非商业性使用-禁止演绎4.0国际许可协议(Attribution-NonCommercial-NoDerivatives 4.0 International)进行授权。若使用本数据集,请引用原论文。 ### 数据集内容 数据集包含用于保护隐私的匿名化字段,结构如下列字段: - postid:数据集内每条帖文的唯一标识符。 - post_text:帖文的文本内容。帖文内的个人可识别信息(Personally Identifiable Information, PII),如提及的账号与URL均已进行哈希处理。 - application_name:发布帖文所用应用或平台名称的哈希值。 - post_language:帖文所用语言。 - in_reply_to_postid:若该帖文为回复帖,则为所回复帖文的匿名化ID。 - in_reply_to_accountid:若该帖文为回复帖,则为所回复账号的匿名化ID。 - post_time:帖文发布时间戳。 - accountid:发布该帖文的账号的唯一匿名化ID。 - account_profile_description:账号持有者在个人主页中填写的简介信息。 - follower_count:数据采集时该账号的粉丝数。 - following_count:数据采集时该账号关注的账号数。 - account_creation_date:账号的创建日期。 - is_repost:布尔型标识,用于指示该帖文是否为转发帖。 - reposted_accountid:若该帖文为转发帖,则为原始发布账号的匿名化ID。 - reposted_postid:若该帖文为转发帖,则为原始帖文的匿名化ID。 - hashtags:帖文内容中包含的话题标签(若有)。 - urls:帖文中分享的URL的哈希值(若有)。 - account_mentions:帖文中提及的账号的匿名化ID(若有)。 - is_control:布尔型标识,用于标记该帖文是否来自对照账号(True表示是,False表示来自信息作战账号)。 不同信息作战行动的数据均存储在本仓库的独立分支中。



