German Traffic Sign Anomaly Detection Dataset
收藏资源简介:
Traffic signs are used to regulate traffic and ensure that it flows smoothly. However, many signs are obscured by dirt, faded paint, scratches, vegetation, stickers, or graffiti, making their meaning unclear. This dataset contains a test set of 369 photos of real traffic signs from 71 different categories. The images were taken with a smartphone camera in various towns in the German state of Schleswig-Holstein and cropped to a square aspect ratio. Additionally, a version of each sign with the background removed is provided. The IDs used to refer to the different sign types originate from the German traffic sign catalogue. Some signs show their normal condition, while others exhibit various anomalies. Since these are real-world examples, some signs have stickers or graffiti with political or other statements. These were not censored to ensure the dataset's realism and utility. Their presence does not imply that the authors approve, support, or agree with these messages. The data is made available for research and development purposes, reflecting real-world conditions. Since classifying traffic signs as anomalous depends on context and the decisions to be derived from the classification, two files provide different binary image-level anomaly classifications. The "corrupt" setting flags all signs as anomalous when something impedes rapid recognition of the sign and its regulatory function. This includes faded colors, heavy pollution or scratches, and occlusion through vegetation, stickers, or graffiti. The "strict" setting considers every sign anomalous if it deviates from the ideal condition by more than a small spot of dirt. The dataset's main application is evaluating methods for detecting anomalies in traffic signs, as done in the related publication "Unified Traffic Sign Anomaly Detection via Template Guidance" by Bischoff & Tomforde (2026), which was accepted at ICMLA 2026. This article uses traffic sign templates as part of the reconstruction-based approach as well as synthetically generated training and validation data. For the purposes of reproducibility and further use, this data is also included here. The synthetic training and validation data comprises only the subset of the 20 signs used in the article, not all 71 sign types in the test set. However, the data generation pipeline is open-source and can be found in the related GitHub repository.



