遇见数据集

Persona Arts Korea

收藏
Zenodo2026-05-11 更新2026-05-26 收录
官方服务:

资源简介:

10,000 synthetic Korean cultural arts personas grounded in the 2024 Survey of Korean Artists (Ministry of Culture, Sports and Tourism). Each persona carries 16 quantitative variables and 17 narrative fields in Korean. Quick facts Personas10,000LanguageKoreanArt fields14Variables per persona16 quantitative + 17 narrativeSource2024 Survey of Korean Artists (population 334,036; respondents 5,059)Source licenceKOGL Type 1Release licenceCC-BY-4.0 Source The dataset is derived from the 2024 Survey of Korean Artists (Korean title: 2024년 예술인 실태조사 통계보고서), released by the Korea Culture and Tourism Institute under Korea Open Government Licence Type 1. KOGL Type 1 permits redistribution and derivative works with attribution. Each quantitative distribution traces to a publicly verifiable cell in the source report. The release bundles the 15 grounding tables (T1–T15) for transparency. Distribution quality All 14 art-field shares match the source population within 0.5 percentage points. All five survey-rate fact checks pass within 1 percentage point. Of 104 age-by-variable cross-tab cells, 3 (2.9 percent) deviate from the source statistics by more than 5 percentage points; the maximum absolute deviation is 5.56 percentage points and the mean absolute deviation is 1.28 percentage points. Intended uses Supervised fine-tuning of Korean language models for persona conditioning in the cultural arts domain Few-shot exemplar pools for prompt engineering involving Korean artist personas Controllable evaluation sets where downstream tasks condition on quantitative anchors with a known joint distribution Out-of-scope uses Substituting surveys of real Korean artists; the personas are synthetic and do not represent individual respondents Distributional reference for Korean culture beyond the registered artist population Privacy-respecting transformation of any individual respondent; no individual records were used in generation Inferring health, political, or other variables not covered by the source survey Limitations The quantitative distributions are grounded in the source statistics, but the narrative fields are generated by a language model conditioned on those quantitative anchors and are not derived from observational data. A persona's lived experience as written is not the experience of any real artist. The released table contains only the quantitative and narrative fields described in the dataset card; no human-rater labels, review statuses, memos, or annotation columns are shipped with the dataset. Companion resources The dataset is also available on Hugging Face Datasets at joonhyungbae/Persona-Arts-Korea.

提供机构:
Zenodo
创建时间:
2026-05-11
二维码
社区交流群
二维码
科研交流群
商业服务