Dataset for: "Cyrillic-MNIST: a Cyrillic version of the MNIST dataset"
收藏资源简介:
Paper under revision. This dataset is related to "Cyrillic-MNIST: a Cyrillic version of the MNIST dataset". It is a new handwritten dataset comprising of raw, grayscale (28x28) and binary(28x28) images of 121,234 samples of 42 Cyrillic letters (33 Russian + 9 Kazakh Ә,Ғ,Қ,Ң,Ө,Ұ,Ү,һ,І) with tabulated file which includes indices for five types of dataset: 1) "filename" - sample file name 2) "letters81" - 81 class with separated lower and uppercase letters excluding those which do not have uppercase versions (ь,ъ,һ). 3) "merged53" - 53 classes with merged lower and uppercase letters except 11 letters which have different writing styles of lower and uppercase. 4) "unbalanced42class" - 42 classes with unbalanced number of images per category both lower and uppercase. 5) "balanced42" - 42 classes with 2000 balanced images per category both lower and upper case. Indices are from 0 to 41, -1 represent omitted samples. 6) "Russian" - 33 classes with 94,217 samples of only Russian Cyrillic handwritten letters. Indices are from 0 to 33, -1 represent omitted samples.



