遇见数据集

拼多多农产品评价中英双语平行语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本数据集为虚拟生成的中英双语平行语料集,模拟拼多多平台农产品评价风格,由百图科技工作室发布(v1.0)。数据集包含10万条语料,每条同时提供中文与英文版本,覆盖正面、中性、负面三种情感倾向,涉及100种农产品品类。中文文本平均16.8字,英文平均10.4词,语料简洁口语化。此外,数据集附有28个细粒度标注维度,涵盖情感强度、用户画像、传播力预测、翻译方式、文化改编等高质量金标准标注,所有内容均通过模板与关键词池组合生成,不含任何真实用户数据。

This is a synthetic Chinese-English parallel corpus that simulates the review style of agricultural products on the Pinduoduo e-commerce platform, released by Baitu Technology Studio (v1.0). The dataset contains 100,000 corpus entries, each with both Chinese and English versions, covering three sentiment polarities: positive, neutral, and negative, and involving 100 categories of agricultural products. The average length of Chinese texts is 16.8 characters, while English texts average 10.4 words, with the corpus featuring a concise and colloquial style. In addition, this dataset is equipped with 28 fine-grained annotation dimensions, including high-quality gold-standard annotations covering sentiment intensity, user profiles, propagation potential prediction, translation modes, cultural adaptation and other related aspects. All content is generated by combining templates and keyword pools, and contains no real user data.

创建时间:
2026-05-26
搜集汇总
数据集介绍
拼多多农产品评价中英双语平行语料集 数据集图片
背景与挑战
背景概述
该数据集是一个由百图科技工作室发布的虚拟生成中英双语平行语料集,包含10万条模拟拼多多平台农产品评价的语料。每条语料均提供中英文版本,覆盖正面、中性和负面三种情感,并附带28个细粒度的高质量标注维度,适用于情感分析、机器翻译等多种自然语言处理任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务