Is The Winner Really The Best? A Critical Analysis Of Common Research Practice In Biomedical Image Analysis Competitions
收藏资源简介:
This data set corresponds to the paper: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions [1] (Experiment: Comprehensive reporting). The key research questions corresponding to this data set were: RQ1: What is the role of challenges for the field of biomedical image analysis (e.g. How many challenges conducted to date? In which fields? For which algorithm categories? Based on which modalities?) RQ2: What is common practice related to challenge design (e.g. choice of metric(s) and ranking methods, number of training/test images, annotation practice etc.)? Are there common standards? RQ3: Does common practice related to challenge reporting allow for reproducibility and adequate interpretation of results? To address these research questions, we aimed to capture all biomedical image analysis challenges that have been conducted up to 2016. To acquire the data, we analyzed the websites hosting/representing biomedical image analysis challenges, namely grand-challenge.org, dreamchallenges.org and kaggle.com as well as websites of main conferences in the field of biomedical image analysis, namely Medical Image Computing and Computer Assisted Intervention (MICCAI), International Symposium on Biomedical Imaging (ISBI), International Society for Optics and Photonics (SPIE) Medical Imaging, Cross Language Evaluation Forum (CLEF), International Conference on Pattern Recognition (ICPR), The American Association of Physicists in Medicine (AAPM), the Single Molecule Localization Microscopy Symposium (SMLMS) and the BioImage Informatics Conference (BII). This yielded a list of 150 challenges with 549 tasks. Next, a tool for instantiating the challenge parameter list introduced in [1] was used by some of the authors (engineers and medical student) to formalize all challenges that met our inclusion criteria as follows: (1) Initially, each challenge was independently formalized by two different observers. (2) The formalization results were automatically compared. In ambiguous cases, when the observers could not agree on the instantiation of a parameter - a third observer was consulted, and a decision was made. When refinements to the parameter list were made, the process was repeated for missing values. Based on the formalized challenge data set, a descriptive statistical analysis was performed to characterize common practice related to challenge design and reporting. [1] Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A. P., Carass, A., Feldmann, C., Frangi, A. F., Full, P. M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B. A., März, K., Maier, O., Maier-Hein, K., Menze, B. H., Müller, H., Neher, P. F., Niessen, W., Rajpoot, N., Sharp, G. C., Sirinukunwattana, K., Speidel, S., Stock, C., Stoyanov, D., Aziz Taha, A., van der Sommen, F., Wang, C.-W., Weber, M.-A., Zheng, G., Jannin, P., Kopp-Schneider, A.: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions. arXiv preprint arXiv:1806.02051 (2018).
本数据集对应论文:《胜者真的最优吗?生物医学图像分析竞赛(biomedical image analysis competitions)中常见研究实践的批判性分析》[1](实验:完整报告)。 对应本数据集的核心研究问题如下: RQ1:生物医学图像分析(biomedical image analysis)领域中,挑战赛(challenge)发挥着何种作用?例如,截至目前已开展多少项挑战赛?其涉及哪些领域?针对哪些算法类别?基于何种成像模态? RQ2:与挑战赛设计相关的常见实践有哪些?例如指标与排序方法的选择、训练/测试图像数量、标注规范等。当前是否存在通用标准? RQ3:挑战赛报告相关的常见实践是否能够保障结果的可复现性与充分解读性? 为解答上述研究问题,我们旨在收录截至2016年已开展的全部生物医学图像分析挑战赛。为获取相关数据,我们对承办或代表生物医学图像分析挑战赛的网站进行了分析,具体包括grand-challenge.org、dreamchallenges.org、kaggle.com,以及生物医学图像分析领域主流学术会议的官方网站:医学图像计算与计算机辅助干预(Medical Image Computing and Computer Assisted Intervention, MICCAI)、国际生物医学成像研讨会(International Symposium on Biomedical Imaging, ISBI)、国际光学与光子学学会(International Society for Optics and Photonics, SPIE)医学影像分会、跨语言评估论坛(Cross Language Evaluation Forum, CLEF)、国际模式识别会议(International Conference on Pattern Recognition, ICPR)、美国医学物理学家协会(American Association of Physicists in Medicine, AAPM)、单分子定位显微镜研讨会(Single Molecule Localization Microscopy Symposium, SMLMS)以及生物图像信息学会议(BioImage Informatics Conference, BII)。经统计,共收录150项挑战赛,包含549个任务。 随后,部分作者(工程师与医学生)使用了文献[1]中提出的挑战赛参数列表实例化工具,将符合纳入标准的所有挑战赛进行标准化整理,流程如下:(1) 最初由两名独立观察员分别对每项挑战赛进行标准化整理;(2) 自动比对两名观察员的整理结果。若出现歧义情形,即两名观察员无法就某一参数的实例化达成一致,则邀请第三名观察员参与决策。若后续对参数列表进行了修订,则针对缺失值重复上述流程。基于标准化整理后的挑战赛数据集,我们开展了描述性统计分析,以刻画挑战赛设计与报告相关的常见实践。 [1] Maier-Hein, L., Eisenmann, M., Reinke, A., Onogur, S., Stankovic, M., Scholz, P., Arbel, T., Bogunovic, H., Bradley, A. P., Carass, A., Feldmann, C., Frangi, A. F., Full, P. M., van Ginneken, B., Hanbury, A., Honauer, K., Kozubek, M., Landman, B. A., März, K., Maier, O., Maier-Hein, K., Menze, B. H., Müller, H., Neher, P. F., Niessen, W., Rajpoot, N., Sharp, G. C., Sirinukunwattana, K., Speidel, S., Stock, C., Stoyanov, D., Aziz Taha, A., van der Sommen, F., Wang, C.-W., Weber, M.-A., Zheng, G., Jannin, P., Kopp-Schneider, A.: Is the winner really the best? A critical analysis of common research practice in biomedical image analysis competitions. arXiv预印本arXiv:1806.02051 (2018)



