结构变化数据集
收藏资源简介:
本研究构建了一个名为“结构变化数据集”的数据集,包含3,000个精心设计的输出空间,共350,000个候选生成文本。该数据集旨在捕捉三种代表性的潜在结构类型:对话行为、情绪和响应结构。这些数据是基于自然发生的对话和指令遵循上下文,但呈现了三种结构类型的可控不确定性。通过分析该数据集,研究发现,在常用的效用函数下,MBR解决方案在不到一半的情况下是簇优的。为了解决这个问题,研究提出了三种新的方法——聚类、结构嵌入和效用截止,以适应效用函数,以考虑候选者在特定结构候选组中的(软)成员资格,同时保持风险最小化的决策理论基础。
This study constructs a dataset named "Structural Variation Dataset", which includes 3,000 meticulously designed output spaces and a total of 350,000 candidate generated texts. This dataset aims to capture three representative latent structural types: dialogue acts, emotions, and response structures. These data are sourced from naturally occurring conversations and instruction-following contexts, while exhibiting controllable uncertainty across the three structural types. Through analysis of this dataset, the study finds that under commonly used utility functions, MBR solutions are cluster-optimal in less than half of the cases. To address this issue, this study proposes three novel methods—clustering, structural embedding, and utility cutoff—adapted for utility functions, which consider the (soft) membership of candidates within specific structural candidate groups while preserving the decision-theoretic foundation of risk minimization.

- 1通过阿姆斯特丹大学, 苏黎世大学, 伦敦大学学院 · 2025年



