DiaBLa
收藏资源简介:
DiaBLa是由爱丁堡大学和LIMSI, CNRS合作创建的一个英法双语书面对话数据集,包含144个自发对话,总计超过5700个句子。该数据集通过众包方式收集,涵盖多种对话主题,并附有细致的人工翻译质量评价。数据集的创建旨在为机器翻译模型评估提供独特资源,并分析机器翻译辅助的通信方式。DiaBLa数据集的应用领域包括机器翻译模型的评估和非正式书面交流中语言行为的分析,旨在解决机器翻译在日常书面交流中的应用问题。
DiaBLa is an English-French bilingual written dialogue dataset co-developed by the University of Edinburgh and LIMSI, CNRS. It contains 144 spontaneous dialogues, totaling over 5,700 utterances. Collected via crowdsourcing, this dataset covers a diverse range of dialogue topics and is accompanied by detailed manual translation quality evaluations. The dataset was created to provide a unique resource for machine translation model evaluation and to analyze machine translation-assisted communication modes. Application scenarios of the DiaBLa dataset include machine translation model evaluation and analysis of linguistic behaviors in informal written communication, aiming to address the practical application challenges of machine translation in daily written communication.




