tyhx/douyin
收藏资源简介:
该数据集包含从抖音(中国版TikTok)收集的1,133,545个独特帖子的元数据,通过雪球采样方法从相关视频推荐中爬取。起始于匹配特定关键词的种子视频,爬虫迭代抓取相关视频以构建多样化的平台内容样本。所有帖子通过aweme_id(唯一视频标识符)进行去重,并按aweme_id的前4位数字分区以高效访问。数据集包含187个字段,核心字段包括视频唯一标识符、描述、创建时间戳、时长、地区代码等,以及互动统计(如点赞数、评论数、分享数、播放数、收藏数)、作者信息(如用户ID、昵称、签名)和内容元数据(如话题标签、音乐、视频文件元数据、位置信息)。数据集仅包含元数据,不包含实际视频或图像内容,且为时间点快照,互动统计反映收集时的状态。
This dataset contains metadata from Douyin short videos, collected by crawling related video recommendations. Starting from seed videos matching specific keywords, the crawler iteratively fetched related videos to build a diverse sample of the platforms content. All posts are deduplicated by aweme_id (unique video identifier) and partitioned by the first 4 digits of aweme_id for efficient access. The dataset includes 187 columns with key fields such as unique video identifier, description, creation timestamp, duration, region code, engagement statistics (e.g., likes, comments, shares, views, saves), author information (e.g., user ID, nickname, bio), and content metadata (e.g., hashtags, music, video file metadata, location data). It contains only metadata without actual video/image content and is a point-in-time snapshot with engagement statistics reflecting collection time.



