CORA, Recipe
收藏资源简介:
本工作构建了两个面向不同领域的偏好数据集,CORA和Recipe,分别用于咖啡店助手对话回应和烹饪食谱程序性描述。CORA包含262条样本,每对包含TTS友好与不友好回应,涵盖符号、缩写、价格、时间等不利元素;Recipe源自RecipeNLG语料库,选取300条食谱,以类似方式构造偏好对。数据集通过自动规则与启发式方法生成,确保文本表面特征的多样性。该数据集旨在对齐大型语言模型,使其原生生成适合语音合成的文本,从而简化TTS流水线、降低延迟,提升多目标平衡下的听感友好性。
This work constructs two preference datasets for distinct domains: CORA and Recipe, which are respectively tailored for coffee shop assistant dialogue responses and procedural descriptions of cooking recipes. CORA consists of 262 samples, with each pair containing TTS-friendly and TTS-unfriendly responses, covering adverse elements such as symbols, abbreviations, prices, time references and other similar factors. Recipe is derived from the RecipeNLG corpus, where 300 recipes are selected and preference pairs are constructed in a similar manner. The datasets are generated through automated rules and heuristic approaches, ensuring diversity in surface-level textual characteristics. These datasets aim to align large language models (LLMs) to natively generate text suitable for speech synthesis, thereby simplifying TTS pipelines, reducing latency, and improving auditory friendliness under multi-objective trade-offs.





