Files attached here are (1) Management journals results (data downloaded from Scopus), (2) Environmental journal results (data downloaded from Scopus), and (3) Scattertext Input.
Text mining has been a dominant approach to extracting useful information from massive unstructured data online. But existing tools for Chinese word segmentation are not ideal for processing social me
This corpus contains the full text of Wikipedia, and it contains 1.9 billion words in more than 4.4 million articles. But this corpus allows you to search Wikipedia in a much more powerful way than is