MIDV-SH: A Synthetic Video Dataset for Robust Laminated Identity Document Recognition and Specular Highlight Detection
收藏资源简介:
Specular highlights significantly degrade the quality of laminated identity document images obtained under uncontrolled conditions, leading to failures in automated data extraction systems. Solving this problem requires not only reliable detection and removal of such artifacts, but also an understanding of how they impact other document recognition tasks, including document localization and field recognition. We present a new synthetic benchmark dataset MIDV-SH, designed for evaluating algorithms for detecting and removing specular highlights from laminated document surfaces, as well as for testing the robustness of document recognition under such challenging conditions. The dataset is generated using physically based rendering with image-based lighting and surface normal maps to simulate realistic document irregularities. It consists of 5,000 videos of 42 types of ID cards in different languages, obtained by rotating the document in the scene to produce highlights with different shapes and intensities. Each frame of the clip has a ground truth illumination map and the same frame without highlights. In addition, for each clip there is a frame-by-frame annotation of the document quad and fields quads with their text values. This allows the benchmark to be used to evaluate the entire document recognition process in video format or the recognition of individual frames.



