传思语源多语种双语对照平行语料
收藏资源简介:
包含100个语种、30个专业领域的海量双语平行语料,可直接用于人工智能自然语言处理的研发和应用、机器翻译引擎研发、大模型研发的训练数据、语音合成训练数据,也可以用作外语教学及研究的翻译记忆库,计算机辅助翻译(CAT)记忆库及高校各专业领域的外文教育教学和研究、大数据管理教育教学及研究等。
A massive parallel bilingual corpus covering 100 languages and 30 professional domains. It can be directly utilized as training data for the research, development and deployment of artificial intelligence (AI) natural language processing (NLP) technologies, the development of machine translation engines, the training of large language models (LLMs), as well as speech synthesis. Additionally, it can serve as a translation memory for foreign language teaching and research, a computer-aided translation (CAT) memory, and support foreign language education, teaching and research across various professional fields in colleges and universities, as well as education, teaching and research related to big data management.




