淘宝网相机评价中英双语平行语料集
收藏资源简介:
本数据集为面向自然语言处理领域的中英双语平行语料集,专门针对淘宝平台相机评价场景进行虚拟生成。数据集包含99998条有效数据,每条数据均包含地道的中文评价文本和对应的英文翻译,并附带丰富的细粒度标注信息,包括情感倾向、情感强度、主题关键词、用户画像、适用场景、传播力评分等高价值维度。数据内容覆盖正面、中性、负面等多种情感类型,涵盖画质、对焦、夜景、便携、性价比等多个产品维度,可用于情感分析、机器翻译、文本分类、风格迁移等多种NLP任务的模型训练与评测。
This dataset is a Chinese-English parallel corpus for natural language processing (NLP), which is synthetically generated specifically for the camera review scenario on the Taobao platform. The dataset consists of 99,998 valid entries, each containing authentic Chinese review texts and their corresponding English translations, along with rich fine-grained annotation information covering high-value dimensions such as sentiment polarity, sentiment intensity, topic keywords, user profiles, applicable scenarios, and virality scores. The data covers various sentiment types including positive, neutral and negative, and spans multiple product dimensions such as image quality, focusing, night photography, portability, and cost-effectiveness. It can be used for model training and evaluation of various NLP tasks such as sentiment analysis, machine translation, text classification, and style transfer.




