CLARA-MeD corpus

DIGITAL.CSIC2022-05-15 更新2026-05-11 收录

下载链接：

https://digital.csic.es/handle/10261/269887

下载链接

链接失效反馈

官方服务：

资源简介：

A collection of 24 298 pairs of professional and simplified texts (>96 million tokens) for automatic medical text simplification in Spanish. A parallel corpus with a subset of 3800 sentence pairs of professional and laymen variants (149 862 tokens) is released as a benchmark for medical text simplification. This dataset was collected in the CLARA-MeD project, with the goal of simplifying medical texts in the Spanish language and reducing the language barrier to patient's informed decision making. In particular, the project aims at developing linguistic resources for automatic medical term simplification in Spanish; and conducting experiments in automatic text simplification.

创建时间：

2022-05-15

5,000+

优质数据集

54 个

任务类型

进入经典数据集