TRENCARD Corpus
收藏资源简介:
TRENCARD Corpus是由巴塞罗那自治大学的Gokhan Dogru创建的一个心脏病学领域的双语语料库,专门用于机器翻译的微调。该数据集包含约800,000个源词和50,000个句子,来源于土耳其心脏病学期刊的双语摘要。数据集的创建过程采用了半自动化的方法,利用翻译工具进行数据质量控制和句子对齐。该数据集主要应用于机器翻译的训练和微调,旨在提高特定领域翻译的准确性和效率。
The TRENCARD Corpus is a cardiology-domain bilingual corpus created by Gokhan Dogru from the Autonomous University of Barcelona, specifically intended for the fine-tuning of machine translation models. This dataset contains approximately 800,000 source words and 50,000 sentences, sourced from bilingual abstracts published in Turkish cardiology journals. The dataset was constructed using a semi-automated approach, leveraging translation tools to perform data quality control and sentence alignment. This corpus is primarily applied to the training and fine-tuning of machine translation systems, with the objective of improving the accuracy and efficiency of domain-specific machine translation.




