遇见数据集

Cantemist corpus: gold standard of oncology clinical cases annotated with CIE-O 3 terminology

收藏
Zenodo2021-06-21 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

<strong>Intro:</strong> Cantemist shared task dataset (divided in train, dev1, dev2 and test). In addition, we include here the Cantemist background set. It contains the train, development and test sets of the three subtasks: cantemist-ner, cantemist-norm and cantemist-coding with Gold Standard annotations. In addition, it contains the documents of the background set, without annotations. <strong>Please cite if you use this dataset:</strong> Miranda-Escalada, A., Farré, E., &amp; Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>. <pre><code>@inproceedings{miranda2020named, title={Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results}, author={Miranda-Escalada, A and Farr{\'e}, E and Krallinger, M}, booktitle={Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings}, year={2020} }</code></pre> <strong>Format:</strong> For subtasks cantemist-norm and cantemist-ner, annotations are distributed in Brat format. See Brat webpage for more information For subtask cantemist-coding, codes are grouped in a TSV file with the following columns (this follows the format used in CodiEsp shared task): filename code <strong>Shared task goal:</strong> In the three subtasks, the goal will be to predict the annotations (either the ANN files or the TSV with the codes) given only the plain text files. <strong>Resources:</strong> <strong>Web</strong> <strong>Citation: </strong>Miranda-Escalada, A., Farré, E., &amp; Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>. <strong>Silver Standard corpus</strong> <strong>Annotation guidelines</strong> <strong>YouTube presentations</strong> <strong>Participant codes</strong> For further information, please visit https://temu.bsc.es/cantemist/ or email us at encargo-pln-life@bsc.es

简介:Cantemist共享任务数据集分为训练集、开发集1、开发集2与测试集。本次发布同时附带Cantemist背景数据集,该数据集涵盖三个子任务(cantemist-ner、cantemist-norm、cantemist-coding)的训练集、开发集与测试集,且所有数据均带有金标准标注(Gold Standard annotations);背景数据集仅包含无标注的原始文档。 使用本数据集请引用: Miranda-Escalada, A., Farré, E., & Krallinger, M. (2020). Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results. In <em>Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings</em>. <pre><code>@inproceedings{miranda2020named, title={Named entity recognition, concept normalization and clinical coding: Overview of the cantemist track for cancer text mining in spanish, corpus, guidelines, methods and results}, author={Miranda-Escalada, A and Farré, E and Krallinger, M}, booktitle={Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2020), CEUR Workshop Proceedings}, year={2020} }</code></pre> 数据格式:针对cantemist-norm与cantemist-ner子任务,标注以Brat格式(Brat)进行分发,详细信息可参考Brat官方网页。针对cantemist-coding子任务,编码结果将整理至TSV文件中,文件包含以下两列(该格式遵循CodiEsp共享任务的规范):文件名、编码。 共享任务目标:在三个子任务中,任务目标均为仅根据纯文本文件,预测得到对应标注(即ANN标注文件或包含编码的TSV文件)。 资源:官方网站、前述参考文献、银标准语料库(Silver Standard corpus)、标注指南、YouTube演示视频、参赛代码。如需进一步信息,请访问 https://temu.bsc.es/cantemist/ 或发送邮件至 encargo-pln-life@bsc.es.

提供机构:
Zenodo
创建时间:
2020-08-10
二维码
社区交流群
二维码
科研交流群
商业服务