CSCD-NS
收藏资源简介:
CSCD-NS是首个专为中文母语者设计的中文拼写检查数据集,由腾讯微信人工智能团队创建。该数据集包含40,000个样本,源自中文社交媒体平台,具有大规模和高质量的特点。创建过程中,研究团队采用了一种新颖的方法,通过模拟输入法输入过程生成伪数据,以更真实地反映实际错误分布。CSCD-NS主要用于提升中文母语者的拼写检查技术,解决现有数据集在规模和错误类型上的不足。
CSCD-NS is the first Chinese spelling check dataset specifically designed for native Chinese speakers, created by the Tencent WeChat AI Team. This dataset contains 40,000 samples sourced from Chinese social media platforms, featuring large scale and high quality. During its development, the research team adopted a novel approach to generate pseudo data by simulating the input process of Chinese input methods, which more authentically reflects the actual distribution of spelling errors. CSCD-NS is primarily used to enhance spelling check technologies for native Chinese speakers, addressing the shortcomings of existing datasets in terms of scale and error types.




