manga109-segmentation
收藏资源简介:
Manga109 Segmentation是一个专注于漫画布局分析和实例分割的纯标注数据集(版本v2.0.0)。该数据集不包含漫画图像,需用户另行获取Manga109图像集。它提供了漫画页面中四种关键元素(文本、拟声词、对话气泡、画格)的实例分割掩码(采用COCO RLE格式),以及这些元素之间的几何包含关系(如文本包含于气泡中)和可用的日文转录文本。标注基于109本漫画书,覆盖10,098个页面,包含454,606个标注实例,并按照书籍不相交原则划分为训练集(87本书,8,128页)、验证集(11本书,1,001页)和测试集(11本书,969页)。v2.0.0版本进行了重大更新:优先使用手动绘制的Zenodo文本掩码,并利用模型在人工标注框内精炼文本/拟声词像素;移除了标注不完整的页面以确保背景质量;所有标注标记为非拥挤状态(iscrowd: 0)。数据集以COCO格式发布,适用于漫画布局分析、实例分割、文本检测和阅读顺序推断等任务。但需注意,部分掩码(尤其是文本/拟声词)为模型辅助生成,气泡/画格掩码继承自上游数据集,且包含关系和阅读顺序字段为几何启发式结果,并非完全精确的人工标注。
Manga109 Segmentation is a purely annotated dataset (version v2.0.0) dedicated to manga layout analysis and instance segmentation. This dataset does not contain any manga images, and users must separately obtain the official Manga109 image collection. It provides instance segmentation masks in COCO RLE format for four key elements in manga pages: text, sound effects, speech bubbles, and panels, along with the geometric inclusion relationships between these elements (e.g., text contained within speech bubbles) and the accompanying Japanese transcriptions. The annotations are based on 109 comic books, covering a total of 10,098 pages and containing 454,606 annotated instances. The dataset is divided into training, validation, and test sets following the book-disjoint principle: the training set includes 87 books with 8,128 pages, the validation set includes 11 books with 1,001 pages, and the test set includes 11 books with 969 pages. Version v2.0.0 includes major updates: prioritizing manually drawn Zenodo text masks, refining text and sound effect pixels within manually annotated bounding boxes using models, removing pages with incomplete annotations to ensure high-quality background annotations, and marking all annotations as non-crowded (iscrowd: 0). The dataset is released in COCO format, and is applicable to tasks such as manga layout analysis, instance segmentation, text detection, and reading order inference. However, it should be noted that some masks (particularly those for text and sound effects) were generated with model assistance; speech bubble and panel masks are inherited from upstream datasets; and the inclusion relationship and reading order fields are geometric heuristic results rather than fully precise manual annotations.




