MultiClinSum-2 Test Dataset: Multilingual Summarization of Clinical Case Reports
收藏资源简介:
Task Overview MultiClinSum-2 is a shared task focused on automatic summarization of clinical case reports across 15 languages. The task challenges participants to develop models capable of condensing lengthy clinical narratives into concise summaries while preserving essential diagnostic and clinically relevant information, and supporting healthcare professionals and researchers in efficiently extracting key clinical insights from biomedical literature. The task is organized by the Barcelona Supercomputing Center's NLP for Biomedical Information Analysis group (NLP4BIA) and promoted by European projects DataTools4Heart and AI4HF. The shared task data combines two complementary sources from the biomedical domain: PMC-Patients Subset: Full case-summary pairs derived from the PMC-Patients subset of PubMed Central, where reference summaries were extracted from specific abstract sections. Originally in English, these cases have been translated into all task languages. Native Case Reports: Case reports selected from PubMed with manually-written summaries by the authors. These cases are natively written in English, Spanish, French, and Portuguese, and were translated into all remaining task languages to ensure comprehensive multilingual coverage. Task Website: https://temu.bsc.es/multiclinsum2/ Test Dataset This repository contains the test dataset for MultiClinSum-2 (multiclinsum2_test_set.zip), comprising 1,000 full clinical case reports per language (15,000 total), derived from the same biomedical sources as the training data: the PMC-Patients subset of PubMed Central and manually curated case reports from PubMed. Each case requires participants to generate a concise summary preserving essential clinical details and information. Languages included: English, Spanish, French, Portuguese, Italian, Russian, Catalan, Norwegian, Danish, Romanian, German, Greek, Dutch, Czech, and Swedish. Participants must submit their generated summaries by May 8, 2026, 15:00 CET. Reference summaries will be released after the submission deadline. For submission guidelines, evaluation details, and task information, visit: https://temu.bsc.es/multiclinsum2/ Important: Reference summaries are NOT included in this release and will be made available after the submission deadline (May 8, 2026, 15:00 CET). Related Datasets Training dataset (https://zenodo.org/records/18887797)Includes ~26,000 full case-summary pairs for each of the 15 task languages. Sample dataset (https://zenodo.org/records/18663291)Includes 50 full case-summary pairs for each of the 15 task languages. Contact For questions about the dataset or shared task, please contact: Miguel Rodríguez-Ortega (miguel.rod.bsc@gmail.com) Eduard Rodríguez-López (edu4bsc@gmail.com) Salvador Lima-López (salvador.limalopez@gmail.com) Martin Krallinger (Krallinger.Martin@gmail.com) License: This work is licensed under a CC BY-NC-SA 4.0 (Creative Commons Attribution 4.0 International) License.
任务概述 MultiClinSum-2是一项聚焦于15种语言临床病例报告自动摘要的共享任务(shared task)。该任务要求参赛者开发能够将冗长的临床叙事文本凝练为简洁摘要的模型,同时保留核心诊断信息与临床相关内容,助力医疗专业人员与研究人员高效从生物医学文献中提取关键临床洞见。 本任务由巴塞罗那超级计算中心生物医学信息分析自然语言处理组(Barcelona Supercomputing Center's NLP for Biomedical Information Analysis group,简称NLP4BIA)主办,并由欧洲项目DataTools4Heart与AI4HF推广。 本次共享任务的数据集融合了生物医学领域的两类互补数据源: PMC-Patients子集:源自PubMed Central(PubMed中央库)的PMC-Patients子集的完整病例-摘要对,其参考摘要从特定摘要章节提取。原始文本为英文,现已被翻译为本次任务涉及的所有语言。 原生病例报告:从PubMed中遴选的、由作者手动撰写摘要的病例报告。此类报告原生语言为英语、西班牙语、法语与葡萄牙语,已被翻译为其余所有任务语言,以实现全面的多语言覆盖。 任务网站:https://temu.bsc.es/multiclinsum2/ 测试数据集 本仓库包含MultiClinSum-2的测试数据集(multiclinsum2_test_set.zip),涵盖每种任务语言的1000份完整临床病例报告(总计15000份),其数据源与训练数据一致:PubMed Central的PMC-Patients子集,以及从PubMed中人工遴选的病例报告。参赛者需为每份病例生成一份保留核心临床细节与信息的简洁摘要。 包含的语言:英语、西班牙语、法语、葡萄牙语、意大利语、俄语、加泰罗尼亚语、挪威语、丹麦语、罗马尼亚语、德语、希腊语、荷兰语、捷克语与瑞典语。 参赛者需于2026年5月8日中欧时间15:00前提交生成的摘要。参考摘要将在提交截止日期后公布。如需了解提交指南、评测细则与任务相关信息,请访问:https://temu.bsc.es/multiclinsum2/ 重要提示:本次发布的数据集未包含参考摘要,参考摘要将在提交截止日期(2026年5月8日中欧时间15:00)后公开。 相关数据集 训练数据集(https://zenodo.org/records/18887797):包含15种任务语言各约26000份完整病例-摘要对。 示例数据集(https://zenodo.org/records/18663291):包含15种任务语言各50份完整病例-摘要对。 联系方式 若您对本数据集或共享任务有疑问,请联系: Miguel Rodríguez-Ortega(米格尔·罗德里格斯-奥尔特加):miguel.rod.bsc@gmail.com Eduard Rodríguez-López(爱德华·罗德里格斯-洛佩斯):edu4bsc@gmail.com Salvador Lima-López(萨尔瓦多·利马-洛佩斯):salvador.limalopez@gmail.com Martin Krallinger(马丁·克拉林格):Krallinger.Martin@gmail.com 许可协议 本作品采用知识共享署名-非商业性使用-相同方式共享4.0国际许可协议(CC BY-NC-SA 4.0,Creative Commons Attribution 4.0 International License)进行许可。



