未提及具体数据集名称
收藏资源简介:
本文未提及具体数据集名称和访问地址,但描述了一种从专利中提取药品制造信息的方法。该方法包括两个主要模型:1)一个用于选择包含制造数据文本片段的方法,2)一个命名实体识别系统,用于提取操作、材料和过程条件的信息。数据集包含208,596个药品相关的专利,通过文本聚类技术、潜在狄利克雷分配(LDA)和k-Means聚类算法识别与制造相关的文本部分。命名实体识别(NER)模型采用深度神经网络,在训练集上取得了84.2%的f1分数。该数据集主要用于解决药品制造领域的信息提取问题,旨在促进新药发现和改善治疗方案。
This paper does not specify the dataset's name and access URL, but presents a method for extracting pharmaceutical manufacturing information from patents. The method incorporates two core models: 1) a model for selecting text fragments containing manufacturing data, and 2) a named entity recognition (NER) system for extracting information on operations, materials and process conditions. The dataset comprises 208,596 pharmaceutical-related patents, and the manufacturing-related text segments are identified via text clustering techniques, Latent Dirichlet Allocation (LDA) and k-Means clustering algorithms. The named entity recognition (NER) model adopts deep neural networks and achieves an F1 score of 84.2% on the training set. This dataset is primarily utilized to address information extraction challenges in the pharmaceutical manufacturing field, aiming to promote new drug discovery and improve treatment regimens.




