遇见数据集

ID-Aaker: a dataset for Indonesian Aaker Brand Personality

收藏
Mendeley Data2026-07-02 收录
官方服务:

资源简介:

this data article introduces ID-Aaker, an Indonesian social media dataset annotated with Aaker’s five brand personality dimensions. The dataset comprises 75,756 rows of Indonesian-language texts collected from X (formerly Twitter) using a custom Python-based web scraping script implemented with Selenium WebDriver. Target accounts were selected based on the alignment of their posting characteristics with one of the five Aaker brand personality dimensions, guided by expert recommendation from a communication researcher. Each account was assigned to a single personality dimension prior to data collection, and all posts from that account were labelled accordingly. The dataset is distributed across five personality classes: competence (19,883 instances), sophistication (17,473 instances), excitement (15,270 instances), sincerity (11,909 instances), and ruggedness (11,221 instances). All data are stored in a single machine-readable CSV file containing the pre-processed text, engagement metrics (favorites, retweets, replies, quotes), word count, language code, and the assigned brand personality label. To protect user privacy, all usernames mentioned within the text were replaced with randomly generated pseudonyms (e.g., @user_4921), and all URLs were replaced with a uniform dummy link format. No personally identifiable information is retained in the published dataset

本数据文章介绍了ID-Aaker数据集,这是一个采用阿克(Aaker)五维品牌个性维度进行标注的印尼语社交媒体数据集。该数据集包含75756条印尼语文本,通过基于Python开发的自定义网络爬虫脚本(基于Selenium WebDriver实现)从X(原Twitter平台)采集所得。目标账号的筛选依据为其发布内容的特征与五大阿克品牌个性维度之一的匹配度,筛选过程由传播学研究者提供专业指导。在数据采集前,每个账号均被分配至单一的个性维度,该账号下的所有帖子均被标注为对应维度。该数据集涵盖五大个性类别:干练型(19883条样本)、精致型(17473条样本)、活力型(15270条样本)、真诚型(11909条样本)与粗犷型(11221条样本)。所有数据存储于单个机器可读的CSV文件中,内容包含预处理后的文本、互动指标(点赞、转发、回复、引用)、词数、语言代码以及分配的品牌个性标注。为保护用户隐私,文本中提及的所有用户名均已替换为随机生成的假名(例如@user_4921),所有URL均替换为统一的虚拟链接格式。发布的数据集中未保留任何可识别个人身份的信息。

创建时间:
2026-05-26
二维码
社区交流群
二维码
科研交流群
商业服务