JobCannon Psychometric Response Dataset
收藏资源简介:
v2 — 44,342 item-level responses across nine instruments and 23 languages. Anonymized, item-level responses to nine open-domain psychometric instruments, collected from real test-takers on JobCannon. Each row is one completed assessment: the raw per-item answers, the computed dimensional scores, and the dominant result type. This is a first-party dataset — our own users' responses, not a re-publication of someone else's data. What is new in v2 5.3× more data. 8,394 → 44,342 item-level responses. Multilingual. 10,535 responses are non-English, across 22 languages besides English. Open psychometric corpora are almost entirely English; this is the part of the release we think is hardest to find elsewhere. A by-locale aggregate layer. Three extra tables give result shares and dimension means per instrument per language, so cross-cultural comparisons do not require re-deriving them from the item-level files. Item-level files (one row per completed assessment): mbti.csv (16-type indicator, 60 items, 14,735 responses), career_match.csv (Mini-RIASEC forced choice, 12 items, 6,506), riasec.csv (Holland Codes, 60 items, 4,947), disc.csv (DISC, 12 items, 4,610), enneagram.csv (Enneagram, 36 items, 3,784), big_five.csv (Five-Factor, 50 items, 3,377), multiple_intelligences.csv (Gardner MI, 40 items, 2,936), eq.csv (Emotional intelligence, 10 items, 2,512), dark_triad.csv (Dark Triad, 18 items, 935). Columns: response_id, locale, year_month, duration_seconds, top_result, score_<dimension>…, q1…qN. Aggregate files (by instrument × language): coverage.csv (198 rows — n, item count, distinct results and median duration per cell, with a publishable flag at n ≥ 100), result_distribution_by_locale.csv (330 rows), dimension_means_by_locale.csv (294 rows). Aggregates suppress small cells: a locale cell needs n ≥ 100 to appear, and a result category with n < 30 is folded into other_below_30. _manifest.json records, per instrument, the raw row count pulled, the count kept after item-count filtering, and the resulting column count, so the numbers quoted here can be checked against the build. Language coverage (item-level responses, all instruments combined): en 29,191 · ja 2,682 · ar 1,533 · ko 1,382 · es 855 · ru 788 · id 660 · fr 608 · pt 414 · uk 314 · th 301 · he 286 · zh 192 · it 144 · de 136 · tr 87 · vi 45 · pl 36 · sv 32 · nl 17 · hi 16 · kk 5 · nb 2. A further 4,616 responses carry no locale and are labelled unknown. Eleven locale labels — ar en es fr id ja ko pt ru th and unknown — reach n ≥ 100 in at least one instrument and therefore appear in the aggregate tables, in 42 instrument × locale cells. Privacy. No personal data is in these files, and none is read in the first place. The build selects exactly six columns — locale, answers, scores, top_result, duration_seconds, created_at — so the identifying columns that exist on the source table (user id, anonymous id, participant name and email, referrer, entry host and path, UTM parameters, cohort and organization ids) cannot reach the output even through a later change to the writer. Two further reductions: response_id is a per-file sequential integer rather than the database id, and the timestamp is truncated to year_month, so a row cannot be joined back to a session by its time. Method. Rows are filtered to each instrument's current item count; a response whose answer array is any other length is a legacy test version or an incomplete and is dropped — 2,767 of 47,109 raw rows (5.9%). score_* columns hold the numeric dimensions exactly as scored by the live site; categorical outputs (MBTI's full type with the identity suffix, for instance) travel in top_result. duration_seconds is blank where the recorded value was ≤ 0 or ≥ 7,200 s. Note on the v1 DOI. 10.5281/zenodo.20686670 is the v1 snapshot — 8,394 responses, not 44,342. Cite this record for v2. Live aggregate norms, recomputed monthly and larger than this frozen snapshot, are published at jobcannon.io/research/distributions. Reliability coefficients for the same instruments are at jobcannon.io/research/reliability.



