全国招投标数据商业挖掘分析数据模型
收藏资源简介:
用TextCNN为base model结合分词技术、招投标领域教据和应用场景生成垂直领域的数据分类模型,基于Bert的base模型,增加相对位置、词性等信息在招投标领域的数据上进行微调人工打标签的方式生成训练集并训练出实体抽取模型Ocr文字识别人工标注图片数据进行paddle-0cr微调生成特定领域的OCR文字识别模型
A vertical-domain data classification model is developed using TextCNN as the base model, combined with word segmentation technology, bidding domain data and application scenarios. For the entity extraction model, the training set is generated through manual annotation, and the BERT-base model is fine-tuned on bidding domain data by incorporating additional information such as relative position and part-of-speech, resulting in the trained entity extraction model. For the domain-specific OCR text recognition model, fine-tuning is performed on paddle-ocr using manually annotated image data.




