APSENet: A Text Line Detection Method of Manchu Archives Based on Instance Segmentation Network
收藏资源简介:
Text line detection is an important link in the digitization of Manchu archives, but there are few relevant studies at present, especially the problem that long text is difficult to detect. This paper proposes a text line detection method of Manchu archives called APSENet based on the idea of PSENet image example segmentation model. This method uses ResNet network to extract the text line features of Manchu archives. By introducing the progressive scale expansion algorithm for the segmentation mask of post-processing network output, it can effectively solve the problem that long text is difficult to detect. By introducing the feature channel attention mechanism, it can solve the problem of large text box margin caused by irrelevant background interference. Experimental results show that the algorithm can achieve good detection results
文本行检测是满文档案数字化进程中的关键环节,但当前相关研究较为匮乏,尤其存在长文本难以检测的痛点。本文基于渐进式尺度扩展网络(PSENet)实例分割模型的思想,提出一种面向满文档案的文本行检测方法APSENet。该方法采用残差网络(ResNet)提取满文档案的文本行特征;通过为后处理网络输出的分割掩码引入渐进式尺度扩展算法,可有效解决长文本难以检测的问题;同时引入特征通道注意力机制,能够解决因无关背景干扰导致的文本框边距过大问题。实验结果表明,该算法可取得良好的检测效果。




