遇见数据集

English and Portuguese CBOW Models from Europarl Corpus, version 7, using FastText with Subwords Option

收藏
Zenodo2022-06-01 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

The models were trained using FastText, model CBOW, 40 epochs, and subwords. Each *.BIN file has its *.VEC file with the vocabulary ordered by frequency. The *.BIN file can return a vector to represent an out-of-vocabulary (OOV) word if the necessary parts of the OOV word were used in training. FastText and Gensim can use these files. The English and Portuguese models are identified in the file name, <strong>_en_</strong> and <strong>_pt_</strong> respectively. An Excel file has the neighborhood changes of some selected words during training on each epoch. A previous exercise to find words with more than one meaning.

提供机构:
Zenodo
创建时间:
2022-06-01
二维码
社区交流群
二维码
科研交流群
商业服务