遇见数据集

Word2Vec Models built from a Collection of French 20th-Century Novels

收藏
Zenodo2022-02-08 更新2026-04-07 收录
数据链接:
官方服务:

资源简介:

The models were trained using the Gensim library for Python, developed by Radim Rehurek, in 2017. All models are based on the same collection of 20th century French novels that covers the period from 1900 to 2010, with a large range of authors and genres respresented. The collection contains approximately 1,200 novels and about 60 million tokens. The models were created using the SGNS (Skip-Gram with Negative Sampling) architecture, the context window was always of size of 6 + 6 around the target word, and the texts were lemmatised and POS-tagged beforehand. POS-Tags remain attached to each token (as in "souris_nom"). Other parameters vary by model: some have 200, some have 300 dimensional vectors; the minimum frequency of the words in the model varies with values of 50, 100 and 200, something which influences the size of the vocabulary and the size of the model.

提供机构:
Schöch, Christof
创建时间:
2022-02-08
二维码
社区交流群
二维码
科研交流群
商业服务