遇见数据集

Datasets for "Exploring the potential of neural machine translation for cross-language clinical NLP resource generation through annotation projection"

收藏
Zenodo2024-09-09 更新2026-06-05 收录
官方服务:

资源简介:

This repository contains the data and additional resources used for the paper "Exploring the potential of neural machine translation for cross-language clinical NLP resource generation through annotation projection". There are four different datasets included, namely: The DisTEMIST, DrugTEMIST and MEDDOPROF corpora in Catalan (DisTEMIST-cat, DrugTEMIST-cat, MEDDOPROF-cat), created through Neural Machine Translation followed by clinical entity mention annotation projection techniques. The projected annotations were further validated manually by bilingual expert annotators, who also provided alternative translations for the annotated terms in case the automatic translation was incorrect. We provide two different versions of the data: (i) The output of the annotation projection process as is, without any further validation. (ii) The validated version of the data fter manual correction by experts. The Catalan Clinical Case Corpus (CataCCC), a collection of 200 clinical case reports in originally written in Catalan covering a variety of clinical specialties. This corpus includes manually validated annotations for diseases, medications and professions created by the same experts who annotated the corpora mentioned above, using the same guidelines and annottaion criteria. License This work is licensed under a Creative Commons Attribution 4.0 International License. Contact If you have any questions or suggestions, please contact us at: - Salvador Lima-López (<salvador [dot] limalopez [at] gmail [dot] com>)- Martin Krallinger (<krallinger [dot] martin [at] gmail [dot] com>)

提供机构:
Zenodo
创建时间:
2024-07-30
二维码
社区交流群
二维码
科研交流群
商业服务