Open Korean Historical Corpus
收藏资源简介:
Open Korean Historical Corpus是一个跨越1300年历史,包含6种语言的开放许可数据集,包括韩国式汉字(Idu)和汉字-韩文混合脚本等代表性不高的书写系统。该语料库包含从7世纪到2025年的1800万份文档和50亿个tokens,来源广泛,从皇家秘书处日记到现代新闻文章。语料库为韩国历史语言学提供了基础资源,并可作为大型语言模型的预训练语料库,以提高其对现代韩文中的汉韩词汇以及古代书写的理解。
The Open Korean Historical Corpus is an open-licensed dataset spanning 1,300 years of history and covering six languages, including less commonly attested writing systems such as Idu (Korean-style Chinese characters) and mixed Hanja-Hangeul scripts. This corpus contains 18 million documents and 5 billion tokens spanning from the 7th century to 2025, with diverse sources ranging from royal secretariat diaries to modern news articles. It serves as a foundational resource for Korean historical linguistics, and can also be utilized as a pre-training corpus for large language models (LLMs) to improve their comprehension of Sino-Korean vocabulary in modern Korean and ancient writing systems.

- 1Open Korean Historical Corpus: A Millennia-Scale Diachronic Collection of Public Domain TextsKAIST, Korea University, New York University, Genentech · 2025年



