BERT fine-tuned CORD-19 NER Dataset

Name: BERT fine-tuned CORD-19 NER Dataset
Creator: Anutariya, Chutiporn; Andres, Frederic; Thant, Shin; Racharak, Teeradaj
License: 暂无描述

IEEE2026-04-17 收录

下载链接：

https://ieee-dataport.org/documents/bert-fine-tuned-cord-19-ner-dataset

下载链接

链接失效反馈

官方服务：

资源简介：

This Named Entities dataset is implemented by employing the widely used Large Language Model (LLM), BERT, on the CORD-19 biomedical literature corpus. By fine-tuning the pre-trained BERT on the CORD-NER dataset, the model gains the ability to comprehend the context and semantics of biomedical named entities. The refined model is then utilized on the CORD-19 to extract more contextually relevant and updated named entities. However, fine-tuning large datasets with LLMs poses a challenge. To counter this, two distinct sampling methodologies are utilized. First, for the NER task on the CORD-19, a Latent Dirichlet Allocation (LDA) topic modeling technique is employed. This maintains the sentence structure while concentrating on related content. Second, a straightforward greedy method is deployed to gather the most informative data of 25 entity types from the CORD-NER dataset.

提供机构：

Anutariya, Chutiporn; Andres, Frederic; Thant, Shin; Racharak, Teeradaj

5,000+

优质数据集

54 个

任务类型

进入经典数据集