遇见数据集

KADR - A Hand-Picked Collection of AI-Generated Soviet Propaganda Posters

收藏
Zenodo2026-07-17 更新2026-08-01 收录
官方服务:

资源简介:

Description This is a hand-picked collection of AI-generated images in the visual language of mid-twentieth-century Soviet propaganda posters. Every image was synthetically generated and then manually curated: from the raw model output I kept only the posters that were stylistically consistent, legible in their category, and free of obvious generation artifacts, so the result is a deliberately assembled dataset rather than a random dump of model output. The collection contains 180 posters, all generated with Nano Banana 2 on the platform Artlist.io over the course of June and July 2026. Methodologically it follows the fruit-SALAD approach (Ohm, Karjus, Tamm & Schich, 2025), which builds controlled, balanced image collections in order to study how image-embedding models perceive similarity; here the same idea is applied to a propaganda corpus. Each image is assembled from clearly defined, controlled ingredients — a background theme and a foreground subject — so that for every image it is known exactly which categories were combined. (An anomaly variant, in which a small out-of-place object from a third category is inserted, exists as an optional bonus and is not part of the official dataset.) Why Soviet propaganda posters? Soviet propaganda posters are a natural testbed for a controlled study of visual similarity because they rely on a compact, highly recognisable vocabulary of motifs — heroic workers, soldiers, farmers, families, orators — recombined within a single, coherent constructivist style. Their strong, stereotyped iconography makes the categories visually distinct and easy to hold apart, while the consistent house style keeps everything else constant. Culturally, these posters were among the most pervasive visual media of the Soviet era: mass-produced and displayed in streets, factories and homes, they were designed to shape everyday perception and collective identity, which makes their motif vocabulary a shared cultural code worth studying in its own right. And because the genre is explicitly persuasive by design, the material connects directly to computational propaganda analysis, where the question of which visual cues a model attends to is exactly what is at stake. Motivation The collection is intended as a controlled benchmark. Because each image carries known ground-truth labels — which category the background belongs to, which category the foreground belongs to, and whether a foreign element was inserted — it can be used to probe how different vision models organise what they see: whether they group posters by their foreground subject (content) or by their background scene (context); whether a deliberately incongruous element, such as a rifle in an otherwise domestic scene, pulls an image toward the "wrong" category in embedding space; and which visual features different models prioritise. These questions mirror the fruit-SALAD design of pitting semantic content against visual style. The optional anomaly bonus set adds a further axis that is relevant to propaganda analysis and explainable AI — testing whether a model follows the intended message or latches onto a spurious, out-of-place cue — but it is not part of the official dataset. Data description The dataset consists of 180 hand-picked posters in JPEG format, at a quality/resolution of 512 px and an aspect ratio of 3:2. All images were produced with Nano Banana 2 through Artlist.io between June and July 2026. Every image is accompanied by metadata recording its combination type, its background category and its foreground category, together with the exact prompt used to generate it. The official dataset contains only coherent (same-category) and mixed (different-category) posters; an anomaly set is provided separately as an optional bonus. Visual style All images share a single, deliberately consistent aesthetic, so that style is held constant and only the controlled category variables change. The house style is that of a vintage 1950s Soviet propaganda poster: socialist realism combined with constructivist composition, a bold and limited palette of crimson red, cream and ochre, flat poster shading with strong dark outlines, a dramatic low-angle heroic composition, and a subtly aged lithograph paper texture. Text, lettering and slogans are deliberately avoided, so that the category signal remains purely visual. Categories and image naming The collection uses six thematic categories, all named in English: Army (soldiers, sailors, officers, parades, fortifications, naval and air power), Agriculture (farming, orchards, dairy, livestock, gardens, harvest), Family (parents and children, domestic and courtyard scenes, everyday life), Industry (factories, steel mills, construction, mining, workers and engineers), Sport (athletes, stadiums, gymnasiums, competitions), and Politics (party congresses, demonstrations, orators, monuments, civic and electoral scenes). Each poster is named using a numeric code, so that the filename alone reveals how the image was built. The name consists of three numbers, in the order background — foreground — index: the first number is the background category, the middle number is the foreground category, and the last number is the running index of the poster within that background–foreground combination. The category codes are: Code Category 0 Industry 1 Family 2 Sport 3 Army 4 Agriculture 5 Politics For example, a poster coded 3–1–2 has an Army background (3) and a Family foreground (1), with the last number giving its index within that combination; a coherent same-category poster simply repeats the category in the first two positions (for instance 2–2–… for a Sport background with a Sport foreground). The optional anomaly bonus set uses an extended version of the same scheme, with a fourth slot for the inserted anomaly. Its filenames consist of four numbers followed by a running index: the first number is the background category, the second is the foreground category, the third is the anomaly category (all using the codes above), and the fourth and fifth positions give the two-digit running index of the poster within that combination. For example, a code of 3–1–2–04 denotes an Army background (3), a Family foreground (1) and a Sport anomaly (2), and is the fourth poster in that combination. Output types The official dataset uses two structural types, defined by how the categories are combined. In coherent images, background and foreground come from the same category (an army foreground on an army background), giving internally consistent posters that act as a control condition. In mixed images, background and foreground come from different categories (an agriculture foreground on an industry background), which tests whether a model attends more to the subject or to the setting. A third, anomaly type — in which the background, the foreground and a small out-of-place detail each come from a different category (for example a family foreground on a sport background with a military rifle inserted as a foreign object) — is provided only as an optional bonus set and is not part of the official dataset. Generation process Originally I intended to build the collection by manually segmenting the main figures from existing, original Soviet posters and compositing them onto different backgrounds, but this proved unworkable — among other reasons because several of the categories I wanted were underrepresented in the available archives, so a balanced grid could not be sourced from real material. I therefore moved to fully synthetic generation, which lets every category be populated evenly and in a consistent style. The final images were produced with Nano Banana 2 from detailed natural-language prompts, each specifying the background category, the foreground category (and, for the optional anomaly bonus, the incongruous object and its position); prompts were designed systematically so that categories are balanced and motif variants rotate, so that images within a group do not strongly resemble or repeat one another. The raw generations were then hand-picked, keeping only images that held the style and showed a recognisable category, and discarding failures and near-duplicates. Known biases and curation notes Generative models reproduce and sometimes amplify the biases of their training data, and this was visible in the raw output. In the Industry category, for example, the model generated male foreground figures far more often than female ones. Because the collection is hand-picked rather than taken as-is, I was able to partly counteract this by deliberately selecting the rarer female industry workers and engineers during curation, in order to bring the category's gender balance closer to even. Curation cannot fully remove model bias, but it makes the final collection more balanced than the raw generations would be. More generally, hand-picking introduces a degree of curator subjectivity — my judgments about what counts as stylistically consistent or clearly in-category shaped the final selection — which is an inherent property of a hand-picked dataset and should be kept in mind for any downstream use. Reproducibility The collection is reproducible. The exact prompts used to generate the images are provided alongside the dataset, as structured CSV files with per-image category labels and as grouped human-readable Markdown. Running these prompts through Nano Banana 2 regenerates content in the same style and category structure. Because generation involves randomness and the model may be updated over time, regenerated images will not be pixel-identical, but they reproduce the same controlled design of style, categories and combination types. The hand-picking step is, by its nature, the one part that reflects individual curation rather than automatic reproduction. Intended use The collection is designed for the controlled evaluation of image-embedding models, following the fruit-SALAD approach; this evaluation has not yet been carried out and is left to future work, so what follows describes the intended procedure rather than results obtained here. Embeddings would be extracted for every image from several vision models (such as CLIP, DINOv2, ViT or ResNet); pairwise similarities would be computed in each embedding space; and one would then measure whether the models organise the images by foreground category or by background category — for instance through a nearest-neighbour self-recognition test and Mahalanobis-distance heatmaps. Using the optional anomaly bonus set, one could additionally test whether an inserted foreign element pulls an image toward the wrong category, and dimensionality reduction (e.g. PCA) could be used to visualise and compare how each model behaves. Ethical note All images are synthetic and depict a historical propaganda aesthetic for research purposes in digital humanities and machine-learning evaluation. They are not endorsements of the ideology they stylistically reference, and the collection is intended for the study of how models perceive and organise propaganda imagery.

提供机构:
Zenodo
创建时间:
2026-07-14
二维码
社区交流群
二维码
科研交流群
商业服务