遇见数据集

StylePTB

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

我们引入了一个大规模的基准 StylePTB,其中 (1) 成对的句子经历了 21 次细粒度的风格变化,跨越了文本的原子词汇、句法、语义和主题转移,以及 (2) 允许建模的多次转移的组合细粒度的风格变化作为更复杂、更高层次的转移的基石。通过在 StylePTB 上对现有方法进行基准测试,我们发现它们难以对细粒度的更改进行建模,并且在组合多种样式时更加困难。

We introduce a large-scale benchmark dataset named StylePTB. It has two core characteristics: (1) Paired sentence pairs undergo 21 fine-grained style changes, covering atomic lexical, syntactic, semantic, and thematic shifts of text; (2) It supports modeling compositional fine-grained style changes through combining multiple transfers, which serve as building blocks for more complex and high-level style shifts. By benchmarking existing state-of-the-art methods on StylePTB, we find that these methods struggle to model fine-grained modifications, and face even greater challenges when combining multiple styles.

提供机构:
OpenDataLab
创建时间:
2022-05-23
搜集汇总
数据集介绍
StylePTB 数据集图片
背景与挑战
背景概述
StylePTB是一个大规模文本风格转换基准数据集,包含成对句子,覆盖21种细粒度的风格变化,涉及词汇、句法、语义和主题层面,旨在评估模型处理复杂风格组合的能力。该数据集由卡内基梅隆大学于2021年发布,用于推动文本风格转换研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务