淘宝网面膜产品评论语料集
收藏资源简介:
本数据集为面向自然语言处理任务的 AI 生成文本语料集,本数据集包含20万条有效评论文本,基于大语言模型通过可控提示词工程批量生成。数据集模拟了淘宝网平台的用户生成内容风格,内容长度严格控制在 2-50 字范围内,涵盖多种情感倾向与用户意图类别。本数据集的核心优势在于包含实体属性标注、用户画像模拟、比较关系识别、购买决策信号、情感强度分级等高价值细粒度标注维度,可广泛应用于情感分析、意图识别、文本分类、用户画像推断、竞品分析、消费决策预测等任务。本数据集全部内容由 AI 生成,不包含任何真实用户数据,不涉及个人隐私。
This dataset is an AI-generated text corpus for natural language processing (NLP) tasks. It contains 200,000 valid review texts, which are batch-generated by large language models (LLMs) via controllable prompt engineering. The dataset simulates the style of user-generated content (UGC) on the Taobao platform, with the text length strictly controlled within 2 to 50 words, covering various sentiment orientations and user intent categories. The core strengths of this dataset lie in its high-value fine-grained annotation dimensions, including entity attribute annotation, user profile simulation, comparative relation recognition, purchase decision signals, and sentiment intensity grading. It can be widely applied to tasks such as sentiment analysis, intent recognition, text classification, user profile inference, competitive product analysis, and consumer decision prediction. All content in this dataset is AI-generated, without any real user data or personal privacy involvement.




