Experimental data for the evaluation of the approach to building PDKG
收藏资源简介:
We evaluate 2 key stages in building a PDKG: Extraction of entity terminology and Extraction of concept using NLP. precision and recall are the most commonly used metrics to measure the performance of extraction algorithms. Precision measures accuracy, the probability of the expected result in all extracted samples, and recall measures the completeness of the extracted result, the probability of the expected result being extracted from the original data. In the evaluation process, a certain number of Chinese patent specification body texts as well as sentences describing design knowledge in the patent were selected for experiments to evaluate the proposed method.
本研究针对专利驱动知识图谱(PDKG)构建的两大核心阶段开展评估:基于自然语言处理(Natural Language Processing,NLP)的实体术语抽取(Entity Terminology Extraction)与概念抽取(Concept Extraction)。 精确率(Precision)与召回率(Recall)是衡量抽取算法性能最常用的两项指标:精确率用于表征抽取结果的准确率,即全部已抽取样本中符合预期结果的概率;召回率则用于表征抽取结果的完整性,即原始数据中符合预期的结果被成功抽取的概率。 本次评估实验选取了一定数量的中文专利说明书正文文本,以及专利中描述设计知识的语句作为实验数据集,以验证所提出方法的有效性。




