ChEMU dataset for information extraction from chemical patents
收藏资源简介:
The discovery of new chemical compounds and their synthesis process is of great importance to the chemical industry. Patent documents contain critical and timely information about newly discovered chemical compounds, providing a rich resource for chemical research in both academia and industry. Chemical patents are often the initial venues where a new chemical compound is disclosed. Only a small proportion of chemical compounds are ever published in journals and these publications can be delayed by up to 3 years after the patent disclosure. In addition, chemical patent documents usually contain unique information, such as reaction steps and experimental conditions for compound synthesis and mode of action. These details are crucial for the understanding of compound prior art, and provide a means for novelty checking and validation. Due to the high volume of chemical patents, approaches that enable automatic information extraction from these patents are in demand. To develop natural language processing methods for large-scale mining of chemical information from patent texts, a corpus is created providing chemical patent snippets and annotated entities and reaction steps.
新型化合物及其合成工艺的发现,对化学工业具有重要意义。专利文献包含了关于新发现化合物的关键且时效性强的信息,为学术界与工业界的化学研究提供了丰富资源。化学专利通常是新型化合物首次公开的渠道。仅有极少数化合物会在期刊上发表,且这类发表往往会比专利公开延迟最多三年。此外,化学专利文献通常包含独特信息,例如化合物合成的反应步骤、实验条件以及作用模式。这些细节对于理解化合物相关现有技术至关重要,同时也为新颖性核查与有效性验证提供了支撑手段。鉴于化学专利的体量庞大,亟需能够从这些专利中自动提取信息的技术方案。为开发可从专利文本中大规模挖掘化学信息的自然语言处理方法,研究人员构建了一个语料库,该语料库包含化学专利片段以及标注后的实体与反应步骤。




