遇见数据集

淘宝网耳机评价中英双语平行语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本数据集为虚拟生成的双语语录语料集,模拟了淘宝网耳机商品评价平台的语录分享内容风格,共包含100,000条高质量中英对照评价数据。每条语录同时提供地道的中文原文和英文翻译,内容长度严格控制在中文2-120字、英文1-80词范围内,涵盖正面、负面、中性等多种情感倾向,以及音质、降噪、佩戴舒适度、续航、性价比、做工品质、连接稳定性等多维度主题类别。所有数据均附有细粒度自动标注,包含语录类型、情感倾向与强度、用户画像模拟、适用场景、传播力评分、翻译质量评估等35+高价值维度,标注置信度为1.0(金标准)。本数据集全部内容为虚拟生成,不包含任何真实用户评价或个人隐私数据,版权归属百图科技工作室,使用无忧。

This dataset is a virtually generated bilingual quote corpus that simulates the style of quote-sharing content on Taobao headphone product review platforms, containing 100,000 high-quality Chinese-English parallel review-style quote entries. Each quote provides both authentic original Chinese text and its corresponding English translation, with strict length limits: 2–120 Chinese characters for the Chinese segment and 1–80 English words for the English segment. It covers diverse sentiment orientations including positive, negative, and neutral, as well as multiple thematic categories such as sound quality, noise cancellation, wearing comfort, battery life, cost-performance ratio, workmanship quality, and connection stability. All data are equipped with fine-grained automated annotations, encompassing over 35 high-value dimensions including quote type, sentiment orientation and intensity, simulated user profiles, applicable scenarios, virality score, translation quality assessment, etc., with an annotation confidence score of 1.0 (gold standard). All content of this dataset is virtually generated, containing no real user reviews or personal privacy data. The copyright is held by Baitu Technology Studio, allowing worry-free utilization.

创建时间:
2026-05-23
搜集汇总
数据集介绍
淘宝网耳机评价中英双语平行语料集 数据集图片
背景与挑战
背景概述
该数据集是一个虚拟生成的双语录语料集,模拟淘宝网耳机商品评价风格,包含约10万条中英对照数据,涵盖多种情感倾向和主题类别,并附带细粒度标注。其适用于自然语言处理领域的多项任务,如情感分析、机器翻译和文本分类等,且无真实用户数据与隐私风险。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务