Spanish Billion Word Corpus and Embeddings主要提供大规模的西班牙语文本数据,包含近15亿词汇,来源于网络上的多种资源。该语料库未进行标注,主要用于语言建模和预训练语言模型等任务。数据以句子为单位,存储在100个文本文件中,并采用CC-BY-SA 4.0协议授权。
This dataset contains the vectors from computing KGloVe embeddings from a Page Rank frequency weighted DBpedia 2016-04 graph. For each entity in the graph, the text file in the zip archive contains
Semantic vectors associated with the paper "Don't count, predict! A systematic comparison of context-counting vs context-predicting semantics vectors" Abstract: context-predicting models (more commo