TYDIP
收藏资源简介:
TYDIP数据集由德克萨斯大学奥斯汀分校计算机科学系创建,专注于九种不同语言的礼貌现象分类。该数据集包含每种语言500个示例,总计4500个示例,每个示例都有三重礼貌标注。数据集内容来源于各自语言的维基百科用户讨论页,旨在捕捉丰富的语言策略。创建过程中,通过精心设计的注释流程确保了注释者之间的高一致性。TYDIP数据集的应用领域包括评估多语言模型、构建礼貌的多语言代理等,旨在解决跨文化交流中的礼貌现象问题。
The TYDIP dataset was developed by the Department of Computer Science at The University of Texas at Austin, focusing on the classification of politeness phenomena across nine distinct languages. Comprising 500 examples per language, the dataset totals 4,500 instances overall, with each entry annotated with tripartite politeness labels. The dataset is sourced from Wikipedia user talk pages in their respective languages, aiming to capture a rich repertoire of linguistic politeness strategies. During its construction, a meticulously designed annotation pipeline was implemented to ensure high inter-annotator agreement. Application scenarios of the TYDIP dataset include multilingual model evaluation and the development of polite multilingual AI agents, among others, with the goal of addressing politeness phenomena in cross-cultural communication.

- 1TyDiP: A Dataset for Politeness Classification in Nine Typologically Diverse Languages德克萨斯大学奥斯汀分校计算机科学系 · 2022年



