遇见数据集

Multilingual Artificial Datasets for AI-Driven Cardiology: the CliTe and ParaCliTe corpora

收藏
Zenodo2026-09-14 更新2026-10-01 收录
官方服务:

资源简介:

Abstract The CliTe corpus comprises the artificial clinical records of 699 original adult patients suffering from heart disease originating from 7 countries in Europe. The texts are written originally in seven languages to mimic the real-world reports found in the respective institutions (see Table 1). ParaCliTe is a subset of CliTe that was manually translated into the remaining languages represented in Table 1. In contrast, ParaCliTe silver is the automatic machine translation of these 699 reports into all non-original languages represented in Table 1, and does not comprise manual revision. CliTe includes emergency department and hospital discharge reports, inpatient and outpatient notes, and referral letters. They have been authored by cardiologists and cardiology researchers between June 2023 and June 2026 within the framework of the DT4H Horizon project (https://www.datatools4heart.eu/). The corpora include the most common heart diseases of adults, such as heart failure, ischemic heart disease and atrial fibrillation, and also less common conditions such as postpartum cardiomyopathy and cardiac amyloidosis (see Table 2). They are also sex balanced and include a wide age-range (19 to 110 years of age). They serve as a valuable resource for training and assessing multilingual clinical NLP techniques and language models, aiding tasks like concept detection and document classification. This work is a collaborative effort involving the professionals from participating hospitals, plus researchers at the NLP4BIA-BSC, Translated (Rome), Vall d’Hebron Research Institute and Amsterdam University Medical Centre. Note: full details can be found in the accompanying journal article associated with this dataset. Background Cardiovascular diseases are the most common cause of death globally. They cause a large number of medical consultations, and AI tools to facilitate the work of specialists are sorely needed. European Union legislation heavily restricts access to real EHR records for research, more so for investigation related Artificial Intelligence (AI). On the other hand, the European Health Data Space aims to find secure solutions to share medical data written in the different local languages of the EU. Under these circumstances, multilingual, shareable non-private clinical texts offer a practical solution. In this context, a European group of clinicians have written artificial clinical reports (i.e., clinical documents based on fictional patients) in their own language (Czech, Dutch, English, Italian, Romanian, Spanish and Swedish). The creation of these artificial reports follows the purposely-made Guiding Principles for Multilingual Generation of Artificial Clinical Reports (https://zenodo.org/records/21220631). These datasets were developed to meet the rising need for such resources and to enable the creation of clinical NLP tools. These tools are tailored to address the unique complexities of real-world clinical language, including dense medical jargon, heavy reliance on abbreviations, typographical errors, misspellings, and unstructured sentence formats. Methods In short, CliTe consists of fictional case reports authored by cardiologists and cardiology researchers from seven major university hospitals throughout Europe (see Table 1) during a three-year period (June 2023 to June 2026). Each of these authors is native in the respective languages shown. This task was guided by experts in the creation of guidelines and biomedical corpora for NLP purposes. Hospital City Language Number CRs Saint Anne’s University Hospital Brno Czech 100 University Medical Centre Utrecht Dutch 100 University College London Hospital London English 100 Fondazione Policlinico A. Gemelli Rome Italian 99 Bucharest Emergency Clinical Hospital Bucharest Romanian 100 Vall d’Hebron University Hospital Barcelona Spanish 100 Karolinska University Hospital Stockholm Swedish 100 Total 699 Table 1: Original clinical sites, locations and languages related to the CliTe corpus. CRs: artificial clinical reports Each of these reports has a unique identifier and the following metadata: language, type of document, sex, age and main heart condition. Data Description CliTe is a collection of 699 fictional clinical records to illustrate presentations of heart conditions in adult patients. The corpus includes a large number of cardiac diseases, with a predominance of ischemic heart disease and heart failure as main conditions as well as additional cardiac and non-cardiac comorbidities, to mimic actual patients presenting to the cardiology department. Most common conditions Conditions <3 fictional patients Ischemic Heart Disease Heart failure Heart valve disease Arrythmia Dilated cardiomyopathy Atrial fibrillation Syncope Lung disease Chest pain Obstructive cardiomyopathy Pericardial disorder Myocarditis Palpitations Cardiac arrest Hypertension Cardiac arrest Congenital heart disease Endocarditis Heart disease (unspecified) Dyspnea Cardiogenic shock Non-ischemic cardiomyopathy Aortic dissection Cardiac amyloidosis Cerebrovascular disease Chemotoxic-induced cardiopathy Dizziness Fatigue Left Ventricular Hypertrophy Postpartum cardiomyopathy Cough Heart murmur Hypercholesterolemia Peripheral artery disease Sepsis Shock Stress Vertigo Other unspecified conditions Table 2 Comprehensive list of diseases represented in CliTe. Each clinical group wrote the fictional accounts as closely as possible to the real reports in their own cardiac departments. As such, they might differ in type of language, use of abbreviations and length of document. ParaCliTe includes five types of medical document: Emergency Department reports; hospital discharge reports; inpatient notes; outpatient notes; and referral letters. Usage Notes The corpora CliTe, ParaCliTe and CliTe silver are intended for use to train and test NLP and AI tools under development. Potential use cases of these corpora are: information extraction from clinical texts (for instance, extraction of diseases, medications and symptoms); report summarisation, to train systems that can effectively produce abridged reports that do not omit essential information; normalisation of entities by means of automatisation of codes (for instance, ICD-10, SNOMED CT). Limitations and non-intended use This corpus is not intended for clinical use. It cannot be reliably used for clinical decision-making, diagnosis, prognosis, treatment planning, or patient care. Any models trained or evaluated using this corpus cannot be referred to as clinically validated, and results obtained from it must not be presented as evidence of clinical performance, effectiveness, or safety. The corpus is to a very limited extent representative of real-world clinical populations or healthcare settings. Consequently, it cannot support claims of generalizability to hospitals, regions, healthcare systems, or routine clinical practice. Likewise, it cannot be used for epidemiological analyses or population-level inference, as its data distributions do not reflect real-world disease prevalence or patient characteristics. In addition, the corpus is unsuitable for longitudinal studies or for evaluating real-world clinical risk, safety, or reliability, including the assessment of rare adverse events, uncommon conditions, or other edge cases. It must not be used as a substitute for real clinical data in the development, validation, or deployment of clinical systems. Finally, the corpus does not capture the operational conditions of real healthcare environments—such as time pressure, clinician workload, interruptions, or workflow constraints—and therefore must not be used to evaluate model performance under realistic clinical operating conditions. Ethics Consent was not required since all patients represented are fictional.

提供机构:
Zenodo
创建时间:
2026-09-08
二维码
社区交流群
二维码
科研交流群
商业服务