遇见数据集

淘宝网相机评价中英双语平行语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本数据集为面向自然语言处理领域的中英双语平行语料集,专门针对淘宝平台相机评价场景进行虚拟生成。数据集包含99998条有效数据,每条数据均包含地道的中文评价文本和对应的英文翻译,并附带丰富的细粒度标注信息,包括情感倾向、情感强度、主题关键词、用户画像、适用场景、传播力评分等高价值维度。数据内容覆盖正面、中性、负面等多种情感类型,涵盖画质、对焦、夜景、便携、性价比等多个产品维度,可用于情感分析、机器翻译、文本分类、风格迁移等多种NLP任务的模型训练与评测。

This dataset is a Chinese-English parallel corpus for natural language processing (NLP), which is synthetically generated specifically for the camera review scenario on the Taobao platform. The dataset consists of 99,998 valid entries, each containing authentic Chinese review texts and their corresponding English translations, along with rich fine-grained annotation information covering high-value dimensions such as sentiment polarity, sentiment intensity, topic keywords, user profiles, applicable scenarios, and virality scores. The data covers various sentiment types including positive, neutral and negative, and spans multiple product dimensions such as image quality, focusing, night photography, portability, and cost-effectiveness. It can be used for model training and evaluation of various NLP tasks such as sentiment analysis, machine translation, text classification, and style transfer.

创建时间:
2026-05-23
搜集汇总
数据集介绍
淘宝网相机评价中英双语平行语料集 数据集图片
背景与挑战
背景概述
该数据集是面向自然语言处理任务的中英双语平行语料集,专门针对淘宝平台的相机评价场景进行虚拟生成,包含约10万条带有细粒度标注的平行语料。它覆盖多种情感类型和产品维度,适用于情感分析、机器翻译、文本分类等多种NLP任务的模型训练与评测。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务