JobCannon Psychometric Response Dataset
收藏资源简介:
v3 — 54,431 item-level responses across nine instruments and 25 languages. Anonymized, item-level responses to nine open-domain psychometric instruments, collected from real test-takers on JobCannon. Each row is one completed assessment: the raw per-item answers, the computed dimensional scores, and the dominant result type. This is a first-party dataset — our own users' responses, not a re-publication of someone else's data. What is new in v3. The corpus grew 22.8% in nine weeks. The part that matters grew more than twice as fast. Non-English responses are up 52.9% — 10,535 → 16,112, against 22.8% for the dataset as a whole. Open psychometric corpora are almost entirely English; this is the part of the release that is hardest to find elsewhere, and it is now 29.6% of every row here. Two languages appear for the first time: Uzbek (uz) and Persian (fa). Hebrew and Ukrainian cross into the aggregate tables. Both had item-level rows in v2 but neither reached the n ≥ 100 floor in any instrument, so neither appeared in a single aggregate cell. Both do now. Publishable aggregate cells: 42 → 57. Thirteen locale labels now reach n ≥ 100 somewhere, up from eleven. Total: 44,342 → 54,431 item-level responses, +10,089. One correction carried over from v2's card: the language table there labelled Ukrainian uk. The files have always used ua, and the table below now matches them — a reader grepping the CSVs for uk found nothing. Item-level files (one row per completed assessment): mbti.csv (16-type indicator, 60 items, 17,761 responses), career_match.csv (Mini-RIASEC forced choice, 12 items, 8,445), riasec.csv (Holland Codes, 60 items, 5,891), disc.csv (DISC, 12 items, 5,444), enneagram.csv (Enneagram, 36 items, 4,412), multiple_intelligences.csv (Gardner MI, 40 items, 4,161), big_five.csv (Five-Factor, 50 items, 3,997), eq.csv (Emotional intelligence, 10 items, 3,100), dark_triad.csv (Dark Triad, 18 items, 1,220). Columns: response_id, locale, year_month, duration_seconds, top_result, score_<dimension>…, q1…qN. Item wording is not distributed — the q columns hold answer values only. Aggregate files (by instrument × language): coverage.csv (224 rows — n, item count, distinct results and median duration per cell, with a publishable flag at n ≥ 100), result_distribution_by_locale.csv (413 rows), dimension_means_by_locale.csv (376 rows). Aggregates suppress small cells: a locale cell needs n ≥ 100 to appear, and a result category with n < 30 is folded into other_below_30. _manifest.json records, per instrument, the raw row count pulled, the count kept after item-count filtering, and the resulting column count, so the numbers quoted here can be checked against the build. Language coverage (item-level responses, all instruments combined): en 33,712 · ja 3,359 · ar 2,656 · ko 2,075 · es 1,471 · id 1,225 · fr 944 · ru 870 · pt 739 · th 546 · ua 506 · he 343 · zh 325 · it 255 · de 212 · tr 187 · pl 103 · vi 93 · sv 56 · fa 55 · nl 34 · hi 23 · uz 14 · nb 11 · kk 10. A further 4,607 responses carry no locale and are labelled unknown. Thirteen locale labels — ar en es fr he id ja ko pt ru th ua and unknown — reach n ≥ 100 in at least one instrument and therefore appear in the aggregate tables, in 57 instrument × locale cells. The unknown bucket is nine rows smaller than in v2 while every instrument grew. Rows are not deleted from the source table; a response that arrived without a locale can acquire one later, which moves it out of unknown and into a named language. So v3 is a superset of v2 by instrument but not cell-for-cell, and the two unknown columns are not directly comparable. Privacy. No personal data is in these files, and none is read in the first place. The build selects exactly six columns — locale, answers, scores, top_result, duration_seconds, created_at — so the identifying columns that exist on the source table (user id, anonymous id, participant name and email, referrer, entry host and path, UTM parameters, cohort and organization ids) cannot reach the output even through a later change to the writer. Two further reductions: response_id is a per-file sequential integer rather than the database id, and the timestamp is truncated to year_month, so a row cannot be joined back to a session by its time. Method. Rows are filtered to each instrument's current item count; a response whose answer array is any other length is a legacy test version or an incomplete and is dropped — 2,804 of 57,235 raw rows (4.9%). score_* columns hold the numeric dimensions exactly as scored by the live site; categorical outputs (MBTI's full type with the identity suffix, for instance) travel in top_result. duration_seconds is blank where the recorded value was ≤ 0 or ≥ 7,200 s. Note on the earlier DOIs. 10.5281/zenodo.20686670 is the v1 snapshot (8,394 responses) and 10.5281/zenodo.21842000 is v2 (44,342). Cite the concept DOI 10.5281/zenodo.20686669 to always resolve to the newest version. One further DataCite handle resolves to this release: Figshare 10.6084/m9.figshare.32668833. The third handle for this dataset, Harvard Dataverse 10.7910/DVN/RL99KS, still resolves to v2 — Harvard's servers stopped accepting requests from our network partway through this release, so the deposit there could not be made. It will be updated once we can reach them again. The dataset is also mirrored on Hugging Face, Kaggle, OSF and GitHub. Live aggregate norms, recomputed monthly and larger than this frozen snapshot, are published at jobcannon.io/research/distributions. Reliability coefficients for the same instruments are at jobcannon.io/research/reliability.



