ClinSpEn Data: Parallel English-Spanish COVID-19 Clinical Cases, Terminology and Ontology Concepts
收藏资源简介:
<strong>ClinSpEn</strong> This repository contains the sample, test and background data for the ClinSpEn track. ClinSpEn is part of the Biomedical WMT 2022 shared task, having the aim to promote the development and evaluation of machine translation systems adapted to the medical domain with three highly relevant sub-tracks: clinical cases, medical controlled vocabularies/ontologies, and clinical terms and entities extracted from medical content. <strong>Data Description</strong> ClinSpEn proposes three different sub-tracks, each based on a different type of clinical data: <strong>- Clinical Cases:</strong> Parallel EN-EN COVID-19 clinical cases. The direction of this sub-track is EN>ES. The dataset’s case reports were carefully selected to cover a wide range of aspects related to the disease: different types of patients (children, adults, elderly and pregnant people, babies), different comorbidities (cancer, mental health issues, immunosuppressed patients) and symptomatology (mild and severe presentations, dermatologic, immunologic and psychiatric manifestations, thrombosis, …). The reports were translated from English to Spanish by a professional medical translator on a first step and revised by a clinical expert on a second step. The sample set files is made up of parallel txt files, with the Spanish version having a “.es” extension and the English files having a “.en” extension. Each report has been parallelized so that every sentence’s line number corresponds to the same sentence’s line number in both languages. The test and background data is made up of a TSV file with three columns: document number, line number and English line. The clinical cases themselves include COVID-19 case reports as well as diverse content extracted from PubMed. <strong>- Clinical Terminology:</strong> Parallel EN-ES clinical terms extracted from medical literature and clinical records, with particular focus on diseases, symptoms, findings, procedures and professions and translated and revised by professional medical translators. The direction of this sub-track is ES>EN. The sample set contains 7 000 terms as a tab-separated file (TSV), with the first column corresponding to English terms and the second column to Spanish terms. The test and background data is made up of a TSV file with two columns: term number and Spanish term. <strong>- Ontology Concepts:</strong> Parallel EN-ES concepts extracted from various open biomedical ontologies and taxonomies and then manually translated by a professional medical translator. The direction of this sub-track is EN>ES. The sample data includes 400 concepts. The terms are presented as tab-separated file (TSV), with the first column corresponding to English terms and the second column to Spanish terms. The third column includes the term’s origin ontology and its correspondent ID, while the fourth one includes a link to the concept in OBO Library. The test and background data is made up of a TSV file with two columns: concept number and English concept. <strong>Related Links:</strong> <strong>- Sub-track website with more information: </strong>https://temu.bsc.es/clinspen/ <strong>- WMT website: </strong>https://www.statmt.org/wmt22/ <strong>- CodaLab: </strong>https://codalab.lisn.upsaclay.fr/competitions/6696/



