遇见数据集

抖音爆款商品评价中英双语平行语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本数据集是由百图科技工作室虚拟生成的中英双语平行语料集,模拟抖音电商平台商品评价的内容风格,共计100,000条有效数据。每条语料同时提供地道的中文和英文版本,中文长度控制在2-120字,英文控制在1-80词。数据涵盖正面(40%)、中性(30%)、负面(30%)三种情感倾向,涉及100款抖音爆款商品类别。所有数据均附有细粒度自动标注,包含情感倾向、情感强度(1-5级)、主题关键词、语料意图、语气风格、传播指数、语言美感评分、翻译方式及双语对齐质量等13个标注字段,可为自然语言处理相关研究提供高质量训练与评测资源。

This dataset is a Chinese-English parallel corpus virtually generated by Baitu Technology Studio, which simulates the content style of product reviews on the Douyin E-commerce platform and contains a total of 100,000 valid entries. Each entry is provided with both authentic Chinese and English versions, where the Chinese text is limited to 2-120 characters and the English text is limited to 1-80 words. The dataset covers three sentiment orientations: positive (40%), neutral (30%), and negative (30%), involving 100 categories of Douyin best-selling products. All entries are equipped with fine-grained automatic annotations, including 13 labeled fields such as sentiment orientation, sentiment intensity (1-5 scale), topic keywords, corpus intent, tone style, propagation index, linguistic beauty score, translation method, and bilingual alignment quality. It can serve as high-quality training and evaluation resources for natural language processing-related research.

创建时间:
2026-05-26
搜集汇总
数据集介绍
抖音爆款商品评价中英双语平行语料集 数据集图片
背景与挑战
背景概述
本数据集是一个由百图科技工作室虚拟生成的、包含10万条有效数据的中英双语平行语料集,模拟抖音电商平台的商品评价内容风格。数据涵盖正面、中性和负面三种情感倾向,涉及100款爆款商品类别,每条语料均附带情感倾向、强度、主题关键词等13个细粒度自动标注字段,适用于自然语言处理相关研究与开发。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务