This dataset contains the subs2vec embeddings for English, as presented in https://zenodo.org/records/17243814. The embeddings were trained on large-scale subtitle corpora and represent semantic vecto
Cleaned text from enwiki-20161001-pages-articles.xml.bz2, saved as enwiki-2016.text.bz2. The data can be used for training word embedding models (e.g., Word2Vec).