Co-occurrence Statistics Data of the Pile and BERT's pretraining data
收藏数据链接:
官方服务:
资源简介:
This repository contains part of data for the paper "Why Do Neural Language Models Still Need Commonsense Knowledge?" This includes the co-occurrence statistics data computed from the Pile dataset and the pretraining data of BERT, saved in the co-occurrence matrix and occurrence matrix in the numpy (.npy) format.
提供机构:
Kang, Cheongwoong创建时间:
2024-06-20



