The Chilean Waiting List Corpus
收藏资源简介:
In this work we describe the Waiting List Corpus consisting of de-identified referrals for several specialty consultations from the waiting list in Chilean public hospitals. A subset of 3000 referrals was manually annotated with 27892 entities, 1272 attributes, and 762 pairs of relations with clinical relevance. <br> The corpus is 68 % medical and 32 % dental. A trained medical doctor or dentist annotated these referrals, and then together with other three researchers, consolidated each of the annotations. The annotated corpus has nested entities, with 35 % of entities embedded in other entities. We use this annotated corpus to obtain preliminary results for Named Entity Recognition (NER). The best results were achieved by using a biLSTM-CRF architecture using word embeddings trained over Spanish Wikipedia together with clinical embeddings computed by the group. NER models applied to this corpus can leverage statistics of diseases and pending procedures within this waiting list. This work constitutes the first annotated corpus using clinical narratives from Chile, and one of the few for the Spanish language. The annotated corpus, the clinical word embeddings, and the annotation guidelines are freely released to the research community. This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.
本研究详述了候诊队列语料库(Waiting List Corpus),其数据源自智利公立医院候诊队列中用于多项专科会诊的去标识化转诊申请。该语料库选取3000份转诊申请作为子集,由人工完成标注,共包含27892个实体、1272项属性以及762组具有临床相关性的关系对。该语料库中68%为医学类文本,32%为牙科类文本。标注工作由经过培训的临床医生或牙科医生完成,随后与另外三名研究人员共同对所有标注结果进行了整合校验。该标注语料库存在嵌套实体结构,其中35%的实体嵌入于其他实体之中。本研究使用该标注语料库开展命名实体识别(Named Entity Recognition, NER)任务并获得初步实验结果,其中最优实验结果由基于双向长短期记忆网络-条件随机场(biLSTM-CRF)架构实现,该架构结合了在西班牙语维基百科上训练得到的词嵌入,以及本团队自研的临床词嵌入。针对该语料库构建的命名实体识别模型,可利用该候诊队列中疾病与待执行诊疗操作的统计特征。本研究构建的标注语料库是首个基于智利临床叙事文本的公开标注语料库,同时也是为数不多的西班牙语临床标注语料库之一。该标注语料库、临床词嵌入以及标注指南均已向研究社区免费公开。本研究成果采用知识共享署名-非商业性使用-相同方式共享4.0国际许可协议(Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License)进行授权。



