Korean embedding files using the different morphological segmentation granularity of the word
收藏数据链接:
官方服务:
资源简介:
Embedding files using the following segmentation: wordUD morphUD +morphUD Based on wordUD there are 9,692,938 sentences and 157,653,628 words (tokenized) including all articles published in The Hankyoreh during 2016 (1.2M sentences), Sejong morphologically analyzed corpus (3M), and Korean Wiki (20201101) (5.3M): ./fasttext skipgram -input input -output embedding -dim 300
提供机构:
Zenodo创建时间:
2022-01-18



