遇见数据集

kiki-bouba-audio-20k

收藏
魔搭社区2026-08-30 更新2026-08-30 收录
官方服务:

资源简介:

# 🔊 Kiki–Bouba, Spoken Aloud (20k Global Responses) ## **Dataset Summary** This dataset is the **audio companion** to [Rapidata/psychology-association-kiki-bouba-etc](https://huggingface.co/datasets/Rapidata/psychology-association-kiki-bouba-etc). In the original dataset, respondents read the question *"Which one is called 'Kiki'?"* as **written text**. Here, respondents instead **hear the word spoken aloud** — the task shows the same two shapes (a rounded blob and a spiky star) while a short audio clip of "kiki" or "bouba" plays as context. The annotator UI requires the clip to finish playing before an answer can be given. The [Bouba–Kiki effect](https://en.wikipedia.org/wiki/Bouba/kiki_effect) is fundamentally a claim about **sound symbolism** — yet most large-scale collections (ours included) presented the word in writing, filtered through each language's orthography. This dataset removes that filter. * **20,000 responses** (~10,000 per word) * **2 spoken words**: "kiki" and "bouba", same speaker ([Waithera Were](https://commons.wikimedia.org/wiki/File:Sw-ke-buba.flac), Wikimedia Commons, CC BY-SA 4.0) * **Same shape images** as the text dataset, pixel-identical * Respondent-level **country, language, age, gender, occupation** metadata * A **hearing check**: an audio validation set ("can you actually hear this?") was attached to the job, and each response carries the respondent's `userScore_audio` reliability score Data collected with the [**Rapidata API**](https://www.rapidata.ai/) in August 2026. Please consider leaving a ❤️ if you find this dataset interesting. *** ## **Noteworthy Findings** ### **1. Hearing the word restores the classic effect that writing reversed** In the text dataset, the global aggregate for *"Which one is called 'Kiki'?"* came out **marginally reversed** — 51.6% chose the **blob** as "Kiki". With the word spoken aloud, the classic direction returns: **55.2%** choose the **spiky** shape for "kiki", and **62.2%** choose the **blob** for "bouba". ![overall](img/overall_congruence.png) Interestingly, the modality shift is not symmetric: "bouba" was *more* strongly associated with the blob when written (72.7%) than when spoken (62.2%), while "kiki" flips from reversed to classic. With audio, the two words land at similar, moderately-congruent levels instead of the text version's asymmetry. ### **2. The Japanese–Arabic contrast survives the modality change** The most striking pattern in the text dataset — **Japanese speakers** overwhelmingly picking the spiky shape as "Kiki" while **Arabic speakers** lean the other way — replicates with spoken audio, so it is not an artifact of how the two scripts render the word. Japanese speakers: **78%** spiky (audio) vs 74% (text). Arabic speakers remain the only large group below chance: **42%** spiky (audio) vs 29% (text) — audio attenuates but does not eliminate the reversal. ![kiki_languages](img/kiki_by_language.png) ### **3. Country extremes: a 29-point spread, with one country inverted** Pooling both words, **58.7%** of responses pick the shape the classic effect predicts. That average hides a wide spread across countries (≥200 responses each): * **Strongest**: Japan **72.9%**, Ecuador 69.3%, Peru 67.1% * **Weakest**: Algeria **44.0%**, Egypt 51.2%, Iraq 52.2% Algeria is the only country whose confidence interval sits **entirely below chance** (44.0%, 95% CI 41–47%) — its respondents actively prefer the *opposite* pairing rather than merely guessing. Iraq and Egypt land at chance, i.e. no detectable association either way. ![countries](img/country_dumbbell.png) The dumbbell also shows **which word carries the effect, and that it differs by country**. Japan is driven by "kiki" (78% spiky vs 67% blob for "bouba"), while several Spanish- and Portuguese-speaking countries show the reverse profile — "bouba" is the stronger of the two. In the countries at or below chance, the deficit is almost entirely on **"kiki"**; their "bouba" responses stay above chance. So the weak aggregate is not a general failure of sound symbolism there — it is specific to one of the two words. | Country | n | Congruent | 95% CI | "kiki"→spiky | "bouba"→blob | |---|---|---|---|---|---| | Japan | 1,716 | **72.9%** | 71–75% | 78% | 67% | | Ecuador | 225 | **69.3%** | 63–75% | 62% | 77% | | Peru | 550 | **67.1%** | 63–71% | 64% | 70% | | Bolivia | 319 | **65.5%** | 60–71% | 64% | 67% | | Portugal | 334 | **64.4%** | 59–69% | 57% | 72% | | France | 404 | **63.1%** | 58–68% | 60% | 66% | | Côte d'Ivoire | 243 | **61.7%** | 55–68% | 58% | 65% | | El Salvador | 213 | **61.0%** | 54–67% | 54% | 68% | | India | 2,297 | **58.0%** | 56–60% | 53% | 63% | | Philippines | 1,801 | **57.7%** | 55–60% | 56% | 59% | | Spain | 1,498 | **55.6%** | 53–58% | 53% | 58% | | Bangladesh | 466 | **55.2%** | 51–60% | 55% | 55% | | Pakistan | 860 | **54.7%** | 51–58% | 52% | 58% | | Iraq | 942 | **52.2%** | 49–55% | 49% | 56% | | Egypt | 2,041 | **51.2%** | 49–53% | 41% | 61% | | Algeria | 912 | **44.0%** | 41–47% | 40% | 47% | ### **4. Age and gender barely matter — and the raw gaps are country composition** The demographic story is mostly a **negative result**, which is itself worth reporting. Demographics are self-reported and optional, so the cells below cover only part of the data (61% of responses carry an age band, 82% a gender; `Other / Unknown / Prefer not to say` is the platform's own third option): | Age | n | Congruent | 95% CI | |---|---|---|---| | 18-29 | 3,962 | 56.9% | 55–58% | | 30-39 | 2,281 | 56.9% | 55–59% | | 40-49 | 2,076 | 59.2% | 57–61% | | 50-64 | 2,368 | 59.5% | 58–61% | | 65+ | 1,482 | 58.3% | 56–61% | | Gender | n | Congruent | 95% CI | |---|---|---|---| | Female | 7,591 | 58.1% | 57–59% | | Male | 4,958 | 60.1% | 59–61% | | Other / Unknown | 3,803 | 60.7% | 59–62% | ![demographics](img/demographics.png) Taken at face value there are two small gaps: respondents aged 50+ are **+2.1pp** more congruent than the 18–29 group, and men **+2.0pp** more congruent than women. Both **disappear once country is held constant** — recomputing each gap *within* country and averaging (weighted by cell size, countries with ≥300 responses) gives **-0.9pp** for age (8 countries) and **-0.1pp** for gender (12 countries). The raw gaps are an artifact of *who answered from where*: the low-congruence countries skew strongly female (65–69% of respondents reporting a gender), while high-congruence Japan and India skew male (41–43% female). Sample composition, not demography, produces the difference — a useful warning for anyone slicing this dataset by demographics without stratifying by country or language. One robustness check in the same spirit: splitting respondents into quartiles by their `userScore_audio` reliability score moves congruence only between 55.9% and 60.4%, so the effect is not carried by inattentive listeners. *** ## **Dataset Structure** One row per spoken word, mirroring the text dataset's schema: * `question`: the instruction shown to respondents * `spoken_word`: `"kiki"` or `"bouba"` * `audio_context`: the audio clip played to the respondent (▶ playable in the dataset viewer) * `option_1` / `option_2`: the two shape images (`option_1` = blob, `option_2` = spiky — same files and same order as the text dataset) * `option_1_selections` / `option_2_selections`: total respondents choosing each option * `audio_license` / `audio_source`: provenance of the audio clip * `detailed_results`: respondent-level list with: * `selection` (`"option_1"` / `"option_2"`), `country`, `language`, `age`, `gender`, `occupation` * `userScore`: Rapidata's global respondent reliability score * `userScore_audio`: reliability score on the audio validation dimension (new vs the text dataset) *** ## **Loading the Dataset** ```python from datasets import load_dataset import pandas as pd ds = load_dataset("Rapidata/kiki-bouba-audio-20k") # respondent-level responses for the "kiki" row kiki = pd.DataFrame(ds["train"][0]["detailed_results"]) print(kiki.groupby("language")["selection"].value_counts(normalize=True)) ``` *** ## **Data Collection** Collected with the **Rapidata API** on a global annotator audience. Differences from the text run: * The word is played as an **audio context** (same speaker for both words); the task cannot be answered before the clip finishes. * Annotators who cannot listen at the moment get an explicit escape and are routed to other tasks, keeping non-listeners out of the data. * An **audio validation set** (a "can you actually hear this?" check) was mixed into sessions — always for new annotators, sampled for established ones — and each response carries the respondent's `userScore_audio`. *** ## **Intended Use** * Sound-symbolism and cross-modal correspondence research * Cross-linguistic / cross-cultural comparison with the paired [text dataset](https://huggingface.co/datasets/Rapidata/psychology-association-kiki-bouba-etc) — same shapes, same platform, different modality * Human perception & semantic association research ## **Limitations** * Single speaker and accent (Kenyan Swahili speaker); accent effects are not controlled * Single-word design without the paired anchor ("this is kiki, that is bouba"), matching the text dataset's phrasing * Language/country composition is not balanced; Arabic speakers are overrepresented * Country and language are heavily confounded (the below-chance and at-chance countries are all Arabic-speaking), so the country differences in finding 3 should not be read as national rather than linguistic effects * Demographic slices are unreliable unless stratified by country — see finding 4 *** ## **Citation** If you use this dataset, please refer back to this page, and consider leaving a like!

提供机构:
maas
创建时间:
2026-08-21
二维码
社区交流群
二维码
科研交流群
商业服务