hk-synthetic-personas
收藏资源简介:
本数据集包含用于竞选调查模拟和语言模型消息预测试的香港居民合成人设。数据主要基于香港2021年人口普查分区概况的汇总分布表生成。数据集核心包含两个部分:一个包含1000个合成人设的广泛池,以及一个从该池中精选的、注重多样性的100人面板。每个人设包含丰富的信息:人口统计字段(如年龄、性别、地区、教育、职业、收入、语言、移民背景等)、模拟准备字段(如媒体习惯、价值观、信任度、可能反对意见、可能接触渠道等)以及可直接用于提示生成的叙事字段(短/长人设卡片)。数据生成流程包括解析官方普查数据、根据加权分布生成人口统计骨架、对竞选相关的小群体进行受控过采样、通过聚类推断12个人口统计细分群体、生成人设卡片、进行自动化验证(如检查重复、覆盖度、多样性)以及最终选择100人面板。该数据集旨在用于探索竞选反应、反对意见、消息风险、渠道差异、语言可及性问题和细分群体层面的反应差异,但明确禁止用于声称真实香港民意、估计投票份额、预测调查结果或替代真实受访者研究。数据集经过验证,确保主要推断群体在精选面板中得到覆盖,且内部具有多样性。使用前建议将合成输出与小型真实香港调查进行对比校准。
This dataset contains synthetic personas of Hong Kong residents for campaign survey simulation and language model message pre-testing. The data is primarily generated based on aggregated distribution tables from the 2021 Hong Kong Population Census by district profiles. The core of the dataset consists of two parts: a broad pool of 1000 synthetic personas and a curated, diversity-focused panel of 100 personas selected from this pool. Each persona includes rich information: demographic fields (such as age, gender, region, education, occupation, income, language, immigration background, etc.), simulation preparation fields (such as media habits, values, trust levels, potential objections, possible contact channels, etc.), and narrative fields (short/long persona cards) that can be directly used for prompt generation. The data generation process involves parsing official census data, generating demographic skeletons based on weighted distributions, controlled oversampling of small groups relevant to campaigns, inferring 12 demographic segments through clustering, generating persona cards, conducting automated validation (e.g., checking for duplicates, coverage, diversity), and finally selecting the 100-person panel. The dataset is intended for exploring campaign responses, objections, message risks, channel differences, language accessibility issues, and response variations at the segment level, but it explicitly prohibits use for claiming real Hong Kong public opinion, estimating vote shares, predicting survey outcomes, or replacing real respondent research. The dataset has been validated to ensure that key inferred groups are covered in the curated panel and that there is internal diversity. It is recommended to compare and calibrate synthetic outputs with small-scale real Hong Kong surveys before use.




