遇见数据集

A Korean raw text collection for creating a language model

收藏
Zenodo2020-12-11 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

<strong>A very large Korean raw text collection for creating a language model </strong> We collected a very large monolingual dataset for Korean, which contains <strong>over 9.6M sentences and 130.6M eojeols</strong>, to create a language model: Korean Wikipedida (https://dumps.wikimedia.org/kowiki/20201101/, 5.3M sentences and 71.8M eojeols, respectively), the Sejong morphologically analyzed corpus (3.0M and 40.0M), and articles from <em>The Hankyoreh</em> daily newspaper during 2016 (1.2M and 18.6M). We preprocessed raw text into morpheme-segmented text using the POS tagging system (park-tyers:2019:LAW). We also attached the POS label to the morpheme-segmented lexicon, and explicitly include a + symbol for consecutive morphemes. 시인/NNG 윤동주/NNP +,/SP 이준익/NNP 감독/NNG 영화/NNG +로/JKB 부활/NNG See https://github.com/jungyeul/sjmorph for the POS tagging system described in park-tyers:2019:LAW.

提供机构:
Jungyeul Park
创建时间:
2020-12-11
二维码
社区交流群
二维码
科研交流群
商业服务