遇见数据集

专用新词发现原子能力

收藏
官方服务:

资源简介:

基于信息熵Entropy、内部凝聚度PMI、闭包子集、停用词和大规模中文词林去噪的一个专用的新词发现能力,解决各个业务场景中面临的未登录词语义缺失带来的模型误差,帮助模型更高效的落地应用。

A dedicated new word discovery capability based on denoising methods that utilize information entropy, pointwise mutual information (PMI) for internal cohesion evaluation, closed subsets, stop words, and a large-scale Chinese lexical repository. This capability addresses model errors caused by semantic deficits of out-of-vocabulary (OOV) terms across diverse business scenarios, facilitating more efficient deployment and practical application of relevant models.

创建时间:
2023-12-12
搜集汇总
数据集介绍
专用新词发现原子能力 数据集图片
背景与挑战
背景概述
该数据集提供专用新词发现能力,通过结合信息熵、内部凝聚度等算法有效识别未登录词,解决语义缺失问题,提升模型应用效果。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务