Chinese HealthNER Corpus是由NYCU NLP Lab收集和标注的医疗命名实体识别数据集。该数据集首先从提供医疗信息的网站、在线健康相关新闻和医疗问答论坛中爬取文章,然后去除所有HTML标签、图像、视频和嵌入的网络广告,并将剩余文本分割成多个句子。数据集包含了10种实体类型,如人体、症状、医疗器材等,并由三名中文专业的本科生进行标注,标注一致性达到84.1%。
The SocialDisNER corpus of the SMM4H 2022 – Task 10 track was manually annotated by medical experts following the SMM4H-SocialDisNER guidelines. These guidelines were adapted from pre
Though Vaccines are instrumental in global health, mitigating infectious diseases and pandemic outbreaks, they can occasionally lead to adverse events (AEs). Recently, Large Language Models (LLMs) hav