遇见数据集

淘宝网面膜产品评论语料集

收藏
国家数据集管理服务平台2026-05-28 更新2026-05-29 收录
官方服务:

资源简介:

本数据集为面向自然语言处理任务的 AI 生成文本语料集,本数据集包含20万条有效评论文本,基于大语言模型通过可控提示词工程批量生成。数据集模拟了淘宝网平台的用户生成内容风格,内容长度严格控制在 2-50 字范围内,涵盖多种情感倾向与用户意图类别。本数据集的核心优势在于包含实体属性标注、用户画像模拟、比较关系识别、购买决策信号、情感强度分级等高价值细粒度标注维度,可广泛应用于情感分析、意图识别、文本分类、用户画像推断、竞品分析、消费决策预测等任务。本数据集全部内容由 AI 生成,不包含任何真实用户数据,不涉及个人隐私。

This dataset is an AI-generated text corpus for natural language processing (NLP) tasks. It contains 200,000 valid review texts, which are batch-generated by large language models (LLMs) via controllable prompt engineering. The dataset simulates the style of user-generated content (UGC) on the Taobao platform, with the text length strictly controlled within 2 to 50 words, covering various sentiment orientations and user intent categories. The core strengths of this dataset lie in its high-value fine-grained annotation dimensions, including entity attribute annotation, user profile simulation, comparative relation recognition, purchase decision signals, and sentiment intensity grading. It can be widely applied to tasks such as sentiment analysis, intent recognition, text classification, user profile inference, competitive product analysis, and consumer decision prediction. All content in this dataset is AI-generated, without any real user data or personal privacy involvement.

创建时间:
2026-05-22
搜集汇总
数据集介绍
淘宝网面膜产品评论语料集 数据集图片
背景与挑战
背景概述
该数据集是一个包含20万条AI生成文本的语料集,模拟了淘宝网面膜产品的用户评论风格,内容长度严格控制在2-50字之间。其核心价值在于提供了实体属性、用户画像、比较关系、购买决策信号和情感强度等多维度细粒度标注,适用于情感分析、意图识别等多种自然语言处理任务,且全部内容为AI生成,不涉及真实用户隐私。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务