Italian GloVe models
收藏资源简介:
Italian GloVe models trained from scratch on a dataset composed of: - <strong>wiki</strong>: a dump of Italian Wikipedia (as of December 15, 2022), comprising 25,548,651 sentences and 526,640,982 words (3.2 GB of raw text);<br> - <strong>webz</strong>: a dataset of Italian news (159,226 documents) from the webz.io platform, crawled in October 2015, containing 44,041,823 sentences and 44,544,385 words (244 MB);<br> - a dataset of 5,510 Italian news articles from the newspaper ModenaToday (<strong>MT</strong>) or 15,115 documents from the Italian version of Reuters (<strong>RCV2</strong>). <strong>glv_wiki_wbz_mt_20_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and MT for 20 epochs <strong>glv_wiki_wbz_mt_50_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and MT for 50 epochs <strong>glv_wiki_wbz_reut_20_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and RCV2 for 20 epochs <strong>glv_wiki_wbz_reut_50_epochs.zip</strong>: GloVe model trained on the dataset consisting of wiki, webz, and RCV2 for 50 epochs
本数据集用于构建从零开始预训练的意大利语全局词向量模型(GloVe),其训练数据集由以下部分构成:<br>- <strong>wiki</strong>:即截至2022年12月15日的意大利语维基百科(Wikipedia)转储数据集,包含25,548,651个句子、526,640,982个单词,原始文本体积为3.2 GB;<br>- <strong>webz</strong>:来自webz.io平台的意大利语新闻数据集,共计159,226份文档,于2015年10月完成爬取,包含44,041,823个句子、44,544,385个单词,原始数据体积为244 MB;<br>- 额外包含两类新闻数据源:其一为来自报纸ModenaToday(简称MT)的5,510篇意大利语新闻文章,其二为意大利版路透社(Reuters,简称RCV2)的15,115份文档。<br><br>本次发布的预训练模型压缩包如下:<br>1. <strong>glv_wiki_wbz_mt_20_epochs.zip</strong>:基于wiki、webz与MT数据集训练20个轮次得到的GloVe模型;<br>2. <strong>glv_wiki_wbz_mt_50_epochs.zip</strong>:基于wiki、webz与MT数据集训练50个轮次得到的GloVe模型;<br>3. <strong>glv_wiki_wbz_reut_20_epochs.zip</strong>:基于wiki、webz与RCV2数据集训练20个轮次得到的GloVe模型;<br>4. <strong>glv_wiki_wbz_reut_50_epochs.zip</strong>:基于wiki、webz与RCV2数据集训练50个轮次得到的GloVe模型。



