Data and Materials for "Can Vision-Language Models Enforce Social Norms? A Replication of the Child Protest Paradigm"
收藏资源简介:
This Zenodo record contains the supplementary materials for the research paper investigating the norm enforcement capabilities of Vision-Language Models (VLMs). The study replicates a developmental psychology paradigm where models observe an action demonstration by a coordinator followed by a norm violation by a puppet. Contents This record contains the following files: pedagogical_stimulus.mp4: Video stimulus showing the coordinator demonstrating the action with pedagogical cues ("Look") followed by the puppet's violation. intentional_stimulus.mp4: Video stimulus showing the coordinator demonstrating the action intentionally ("Hmm") followed by the puppet's violation. accidental_stimulus.mp4: Video stimulus showing the coordinator demonstrating the action accidentally ("Oops!") followed by the puppet's violation. vlm_responses_classified.csv: The complete dataset of VLM text outputs (2700 responses) and their classifications. system_prompt.txt: The system prompt used to configure the VLMs. Stimuli Files (*_stimulus.mp4) These video files were presented to the VLMs and were generated using Veo 3. Action Description: Coordinator's Action: Pushes a multicolored wooden cube horizontally across a table using a wooden spoon. The manner varies based on the condition (pedagogical, intentional, accidental). Puppet's Action (Violation): Balances the wooden spoon on top of the wooden cube and pushes both items across the table using its hands. Model Responses File (vlm_responses_classified.csv) This CSV file contains the complete set of 2700 responses generated by the VLMs across all conditions. Columns: Video_Name: The type of coordinator demonstration video used (pedagogical, intentional, accidental). Model: The VLM used (gemini-2.5-flash, gemini-robotics-er-1.5-preview, gemini-flash-lite-latest). Temperature: The temperature setting used for generation (0.0, 0.7, 1.0). Trial: The trial number for the specific combination (1-100). Model_Response: The full text output generated by the VLM. Classification (1=protest, 0=no_protest): The manual classification of the response (1 = Protest, 0 = No Protest). System Prompt File (system_prompt.txt) This plain text file contains the exact system prompt provided to each VLM before presenting the video stimulus. This prompt was designed based on the Belief-Desire-Intention (BDI) framework to encourage reflective and natural reactions. Protest Coding Definition Responses were coded as Protest (1) if the VLM's output included any intervention referencing the coordinator's demonstrated action as the standard against which the puppet's action was evaluated. This included: Direct corrections of the puppet. Questions directed to the coordinator about the puppet's deviation. Expressions of confusion or surprise specifically highlighting the difference in actions. Suggestions or demonstrations of the "correct" way to perform the action. Responses were coded as No Protest (0) if they were purely descriptive, narrative, praised the puppet's action without referencing the coordinator's standard, or did not otherwise indicate a normative judgment based on the demonstration.



