电商平台粉底液产品评论语料集
收藏资源简介:
粉底液AI合成双模式文本语料集是一个面向美妆垂直领域的大规模 synthetic corpus,包含中文和英文两个子集,各 50 万条,总计 100 万条数据。每条数据以电商评论为载体,配套纯文本、元数据(平台/用户/产品画像)和标注数据(情感极性/意图/实体/比较关系)三种文件格式。标注维度涵盖 20 余个字段,包括情感倾向与强度、混合情感标记、正面/负面评价面、评论意图、实体属性(品牌/色号/质地/功效)、用户画像(肤色/肤质/年龄/预算)、品牌比较关系、购买意向与复购信号等。数据模拟覆盖淘宝、京东、拼多多、抖音、小红书、Amazon、eBay 等主流电商平台,发布时间跨度 2021—2024 年,适用于需要中英双语、多维度标注的 NLP 研究与工程场景。
The AI-Synthesized Dual-Mode Text Corpus for Foundation Makeup is a large-scale synthetic corpus targeting the vertical beauty and cosmetics domain. It comprises two subsets in Chinese and English, each containing 500,000 entries, totaling 1 million data entries in all. Each entry takes e-commerce customer reviews as the carrier, and is accompanied by three file formats: plain text, metadata (platform, user and product profiles), and annotated data (sentiment polarity, intent, entities and comparison relationships). The annotation dimensions cover more than 20 fields, including sentiment orientation and intensity, mixed sentiment markers, positive/negative review aspects, review intent, entity attributes (brand, shade, texture, efficacy), user profiles (skin tone, skin type, age, budget), brand comparison relationships, purchase intention and repurchase signals, among others. The simulated data covers mainstream e-commerce platforms such as Taobao, JD.com, Pinduoduo, Douyin, Xiaohongshu, Amazon, eBay, etc., with a release period spanning from 2021 to 2024. This corpus is suitable for NLP research and engineering scenarios requiring Chinese-English bilingual and multi-dimensionally annotated data.




