遇见数据集

拼多多家居用品评价中英双语平行语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

拼多多家居用品评价中英双语平行语料集 v1.0 是一个由算法虚拟生成的大规模双语语料库,包含 100,000 条中英双语平行语料。该数据集模拟了拼多多平台的家居用品评价内容风格,每条语料同时提供地道的中文和自然流畅的英文版本。内容长度严格控制在中文 2-120 字、英文 1-80 词范围内,涵盖 100 种家居日用产品类别,情感分布为正面 40%、中性 30%、负面 30%。所有数据均为合成生成,不包含任何真实语料或个人隐私信息,版权清晰,使用无忧。数据集还附有细粒度自动标注,包含语料类型、适用场景、作者/来源模拟、情感强度、语言风格等高价值维度,可作为情感分析、语料分类、机器翻译、文本风格迁移等任务的训练与评测资源。

Pinduoduo Home Goods Reviews Bilingual Parallel Corpus v1.0 is a large-scale bilingual corpus generated algorithmically, consisting of 100,000 Chinese-English parallel text pairs. This dataset simulates the writing style of home goods reviews on the Pinduoduo platform, with each entry providing both authentic Chinese and naturally fluent English versions. The length of each text pair is strictly constrained: 2 to 120 Chinese characters for the Chinese segment, and 1 to 80 words for the English segment. It covers 100 categories of household daily products, with an emotional distribution of 40% positive, 30% neutral, and 30% negative. All data is synthetically generated, containing no real textual materials or personal private information, with clear copyright and hassle-free usage. The corpus is additionally equipped with fine-grained automatic annotations covering high-value dimensions including text type, applicable scenarios, simulated author/source, emotional intensity, and linguistic style. It can serve as a training and evaluation resource for tasks such as sentiment analysis, text classification, machine translation, and text style transfer.

创建时间:
2026-05-26
搜集汇总
数据集介绍
拼多多家居用品评价中英双语平行语料集 数据集图片
背景与挑战
背景概述
该数据集是一个由算法虚拟生成的大规模中英双语平行语料库,包含10万条模拟拼多多平台家居用品评价风格的语料,涵盖100种产品类别并具有特定的情感分布。它提供了细粒度的自动标注,适用于情感分析、机器翻译等多种自然语言处理任务的训练与评测。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务