遇见数据集

CyberCan lexicon

收藏
NIAID Data Ecosystem2026-03-14 收录
数据链接:
官方服务:

资源简介:

Text mining has been a dominant approach to extracting useful information from massive unstructured data online. But existing tools for Chinese word segmentation are not ideal for processing social media text data in Cantonese. This project developed CyberCan, a lexicon of contemporary Cantonese based on more than 100 million pieces of internet texts. The details regarding the creation of the lexicon could be found here: https://osf.io/preprints/socarxiv/tyjr7

创建时间:
2022-09-16
二维码
社区交流群
二维码
科研交流群
商业服务