BOUQuET
收藏资源简介:
BOUQuET是一个多中心、多语域的数据集和基准,由FAIR at Meta等机构创建。该数据集专门设计为非英语中心,包含23种语言,旨在服务于多语言翻译的准确性。数据集内容手工制作,涵盖多种语言特点,并以段落形式组织,超越句子级别。BOUQuET涵盖了广泛的领域,适用于多种语言处理任务,特别适合开放倡议,可扩展至任何书面语言的多向平行语料库。
BOUQuET is a multi-center, multi-register dataset and benchmark developed by organizations including FAIR at Meta and other research institutions. This dataset is specifically designed with a non-English-centric orientation, encompassing 23 languages, with the goal of advancing the accuracy of multilingual translation. The content of the dataset is manually curated, covers a wide range of linguistic characteristics, and is structured at the paragraph level, exceeding sentence-level boundaries. It spans diverse domains and is applicable to multiple natural language processing tasks. It is particularly well-suited for open collaborative initiatives, and can be extended as a multi-directional parallel corpus for any written language.

- 1BOUQuET: dataset, Benchmark and Open initiative for Universal Quality Evaluation in TranslationFAIR at Meta, University College London, University of the Basque Country (UPV/EHU) · 2025年



