CzechDocs
收藏资源简介:
CzechDocs是由查理大学团队构建的一个多语言并行格式化文档数据集,旨在支持评估在翻译过程中保留文档格式的机器翻译系统。该数据集包含316个文档的并行语言变体,涵盖HTML、DOCX和PDF三种格式,总计60,153个可翻译片段、271,111个单词和126,833个标记,数据主要来源于捷克政府网站、教育门户和公共服务平台等真实场景。其创建过程涉及网络爬取、手动对齐、格式转换和预处理流水线,确保了片段级对齐和标记的高密度保留。该数据集专注于解决本地化内容中标记感知翻译的评估难题,特别适用于危机沟通、公民融合材料等多语言内容快速部署领域,为格式保持机器翻译的研究提供了关键资源。
CzechDocs is a multilingual parallel formatted document dataset constructed by the team from Charles University, which aims to support the evaluation of machine translation systems that preserve document formatting during translation. This dataset contains parallel language variants of 316 documents, covering three formats: HTML, DOCX and PDF, with a total of 60,153 translatable segments, 271,111 words and 126,833 tokens. The data is mainly sourced from real-world scenarios such as Czech government websites, educational portals and public service platforms. Its creation process involves web crawling, manual alignment, format conversion and a preprocessing pipeline, ensuring segment-level alignment and high-density retention of tokens. This dataset focuses on addressing the evaluation challenge of token-aware translation in localized content, and is particularly suitable for fields requiring rapid deployment of multilingual content such as crisis communication and citizen integration materials, providing a critical resource for research on format-preserving machine translation.





