The data set consists of 12 subsets(csv)and one Readme file(txt) in total, which are Chinese and Tibetan new words data and usage instructions from 2017 to 2022. The data item of Chinese new words con
A compressed folder containing 31075 numerical vectors. Each one represents word frequencies of an EBook from Project Gutenberg written in English. Vectors are named containing the ID number of their