SPAMID-PAIR
收藏Mendeley Data2024-03-27 更新2024-06-26 收录
下载链接:
https://data.mendeley.com/datasets/fj5pbdf95t
下载链接
链接失效反馈官方服务:
资源简介:
Data post-comment pairs were collected from 13 selected Indonesian public figures (artists) / public accounts with more than 15 million followers and categorized as famous artists. It was collected from Instagram using an online tool and Selenium. Two persons labeled all pair data as an expert in a total of 72874 data. The data contains Unicode text (UTF-8) and emojis scrapped in posts and comments without account profile information. It contains several fields: -igid: Account ID, -comment: Comment of a post, -post: Post from an ID, -emoji: Whether the data contains emojis or not (1 or 0), -spam: Whether the data is spam or not (1 or 0), -lengthcomment: The character length of the comment, -lengthpost: The character length of the post, -countemojicomment: Number of emoji symbol characters in comments, -countemojicommentuniq: Number of emoji symbol characters in comments (unique), -countemojipost: Number of emoji symbol characters in posts, -countemojipostuniq: Number of emoji symbol characters in the post (unique)
本数据集的评论-帖子配对样本采集自13位经筛选的印度尼西亚公众人物(艺人)/粉丝量超1500万的知名公众账号,且均归类为知名艺人。数据通过在线工具与Selenium从Instagram平台采集完成。共计72874条配对数据由两名作为领域专家的标注人员完成标注。
本数据集包含以UTF-8编码的Unicode文本与从帖子及评论中爬取的表情符号,未包含账号个人资料信息。数据集包含以下字段:
- igid:账号ID
- comment:帖子评论内容
- post:对应账号发布的帖子内容
- emoji:数据是否包含表情符号(1代表包含,0代表不包含)
- spam:数据是否为垃圾内容(1代表是垃圾内容,0代表非垃圾内容)
- lengthcomment:评论的字符长度
- lengthpost:帖子的字符长度
- countemojicomment:评论中的表情符号字符总数
- countemojicommentuniq:评论中的唯一表情符号字符数
- countemojipost:帖子中的表情符号字符总数
- countemojipostuniq:帖子中的唯一表情符号字符数
创建时间:
2024-01-23



