JobCannon Entertainment Quiz Response Dataset
收藏资源简介:
v1. 77,284 item-level responses to five for-fun quizzes, in 24 languages. Anonymized, item-level answers to five entertainment quizzes taken by real visitors on JobCannon between March and August 2026. One row is one completed quiz: the raw per-item answers, the category totals the site computed from them, and the result the taker was shown. We collected all of it on our own traffic. It is not a repackaging of somebody else's file. Read this part before you use it. These five quizzes are not psychometric instruments. Nobody validated them. There is no norm sample behind them and no reliability coefficient, and they were never built to measure any particular construct. "Spirit animal" is not a trait. The result a taker sees is the largest of a handful of hand-authored category counters, and those categories were picked because they make a fun result page. So please don't read a mean here as a population estimate, and there is no published norm for any of the five to compare these numbers against. The useful thing in the file is the behaviour around the answers: how people respond to forced-choice items in twenty-four languages, how long they take, and how answer patterns and result shares move between language audiences looking at identical items. The sample is big enough and clean enough for that kind of question. If you want scored instruments with published norms behind them, that is a separate release: the JobCannon Psychometric Response Dataset, nine instruments, no overlap with this one. Both come out of the same generator, under the same privacy rules and the same suppression thresholds, and we kept the two test lists apart on purpose, so any given row appears in exactly one of them. Files. Five item-level tables, one row per completed quiz: Jungian archetype (24 items, 12 result categories, 31,081 responses), Spirit animal (12 items, 8 categories, 17,520), Mental age (12 items, 5 bands, 10,985), Aura colour (10 items, 7 categories, 10,929), Past life era (12 items, 8 categories, 6,769). Columns are response_id, locale, year_month, duration_seconds, top_result, one score_ column per category, then q1 to qN. Item wording is not distributed; the q columns hold answer values only. Three aggregate tables sit beside them: coverage (126 rows), result_distribution_by_locale (233 rows) and dimension_means_by_locale (290 rows), each broken out by quiz and language. Small cells are suppressed. A language cell needs n of at least 100 to appear, and inside a cell any result category under n of 30 is folded into other_below_30. 37 of the 121 quiz by language cells clear that floor; the other 84 are present with publishable=false so you can see they exist without anyone computing a share off eleven people. Language coverage. English is 46.0% of the rows, so most of this file is in some other language: 41,743 rows, of which 40,236 carry a named language. Japanese alone is 26,769 rows, around eight times the Japanese sample in our psychometric release. That mass is bunched into one quiz, which limits what you can do with it: 22,223 of the Japanese rows sit in the Jungian archetype quiz, where Japanese takers are 71.5% of everyone who finished. Arabic behaves differently, and past_life and aura_color go the other way at 86% and 78% English. Check the coverage table for the cell you care about before you assume a language is represented in the quiz you are looking at. A further 1,507 rows arrived without a locale and are labelled unknown; treat those as missing values. Privacy. No personal data is in these files, and none is read in the first place. The build selects exactly six columns (locale, answers, scores, top_result, duration_seconds, created_at). Everything identifying on the source table is therefore out of reach of the output even if somebody later changes the writer: user id, anonymous id, participant name and email, referrer, entry host and path, UTM parameters, cohort id and organization id. On top of that, response_id is a per-file sequential integer instead of the database id, and the timestamp is cut back to year_month, so no row can be joined to a session by its time. Method. Rows are filtered to each quiz's current item count; a response whose answer array is any other length is either a legacy version of the quiz or an incomplete one, and it is dropped, which is 2,649 of 79,933 raw rows, 3.3%. The score_ columns hold the category totals exactly as the live site computed them, with nothing recomputed for the release. duration_seconds is blank where the recorded value was at or below 0 or at or above 7,200 s, roughly 2% of rows. Nothing is weighted, balanced or down-sampled: language shares are traffic shares, and traffic follows wherever the quiz got shared. License. CC-BY-4.0. If you publish anything off it, please cite the DOI.



