遇见数据集

Decade-level Word2Vec models from automatically transcribed 19th-century newspapers digitised by the British Library (1800-1919)

收藏
Zenodo2023-05-24 更新2026-05-26 收录
数据链接:
官方服务:

资源简介:

Word embeddings trained on a 4.2-billion-word corpus of 19th-century British newspapers using Word2Vec and the following parameters: <pre><code>sg = True min_count = 5 window = 5 vector_size = 100 epochs = 5</code></pre> The embeddings are divided into periods of ten years each. Unlike those in this repository, these were not aligned and OCR errors skimmed from the vocabulary. See related GitHub repository for the full documentation: https://github.com/Living-with-machines/DiachronicEmb-BigHistData Project website (Living with Machines): https://livingwithmachines.ac.uk/

提供机构:
Zenodo
创建时间:
2023-05-02
二维码
社区交流群
二维码
科研交流群
商业服务