The Sogou corpus is a widely used Chinese text classification dataset, sourced from Sogou News, containing 17,910 text samples. It covers multiple news categories, providing comprehensive support for
This corpus contains the full text of Wikipedia, and it contains 1.9 billion words in more than 4.4 million articles. But this corpus allows you to search Wikipedia in a much more powerful way than is