遇见数据集

拼多多平价服饰评价中英双语平行语料集

收藏
国家数据集管理服务平台2026-05-29 更新2026-05-30 收录
官方服务:

资源简介:

本数据集为面向自然语言处理(NLP)研究的中英双语平行语料集,模拟了【拼多多】平台上平价服饰评价的内容风格,共包含100000条有效语料。每条语料包含中文评价文本及其地道英文对应翻译,内容长度严格控制在中文2-120字、英文1-80词范围内。数据集覆盖了T恤、卫衣、牛仔裤、连衣裙、衬衫、针织衫、外套、阔腿裤等多种服饰品类,涵盖正面好评、负面差评、中性反馈等多元情感倾向,以及简约、韩系、复古、通勤、文艺等多种穿搭风格。全部内容由大语言模型虚拟生成,不涉及任何真实用户数据。

This Chinese-English parallel corpus dataset is intended for natural language processing (NLP) research. It simulates the content style of affordable apparel reviews on the Pinduoduo platform, and contains a total of 100,000 valid corpus entries. Each entry includes a Chinese review text and its natural, idiomatic English translation, with the length strictly controlled: 2 to 120 Chinese characters for the Chinese content and 1 to 80 English words for the translated text. The dataset covers a wide range of apparel categories, including T-shirts, hoodies, jeans, dresses, shirts, knitwear, jackets, wide-leg pants and more. It encompasses diverse emotional tendencies such as positive reviews, negative reviews and neutral feedback, as well as various dressing styles including minimalist, Korean-style, vintage, commuter-style and literary/artistic styles. All content in this dataset is virtually generated by large language models (LLMs) and does not involve any real user data.

创建时间:
2026-05-26
搜集汇总
数据集介绍
拼多多平价服饰评价中英双语平行语料集 数据集图片
背景与挑战
背景概述
该数据集是一个面向自然语言处理研究的中英双语平行语料集,模拟了拼多多平台上的平价服饰评价风格,包含10万条虚拟生成的中文评价及其地道英文翻译。它覆盖了多种服饰品类、情感倾向和穿搭风格,适用于情感分析、机器翻译等NLP任务。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务