apol/spain-reference-personas-frontier
收藏资源简介:
Spain Reference Personas Frontier 是一个开放的合成参考人口和基准数据集,专门为西班牙设计,用于评估和设计具有社会基础的人工智能系统。该数据集并非观测到的微观数据、调查、对真实公民的预测,也不能替代实地调查、行政数据或领域特定验证。它旨在支持在西班牙运行的AI系统的模拟、评估、提示条件设置、子组分析、服务设计、本地化评估和基准开发。数据集采用分层架构,包括合成人口层(如persona_core和household_core,提供稳定的成人人物结构、家庭链接、经济和地域背景)、LLM交互层(如persona_views和actor_state_init,提供紧凑的提示视图和可变的模拟状态支架)和评估层(如benchmark_tasks,包含可重复的任务、保留分割和验证指标)。数据集包含1,000,000个成人人物、536,741个家庭、6,350,524个LLM面向视图、1,000,000个行动者状态行和1,800个基准任务,总行数8,889,089,大小5.584 GB。它适用于可控模拟、评估、服务设计、研究原型和基准开发,但不适用于监视、移民执法、个人层面说服、政治微目标定位、选民操纵、资格决策、信贷、就业、住房决策或关于真实公民的主张。数据集还提供了详细的统计信息、评估指标(如区域和年龄校准误差)和效率数据(如视图层令牌使用情况),并强调负责任使用,确保合成人物不被视为真实个体。
Spain Reference Personas Frontier is an open synthetic reference population and benchmark substrate for evaluating and designing socially grounded AI systems for Spain. It is not observed microdata, not a survey, not a prediction of real citizens, and not a substitute for fieldwork, administrative data, or domain-specific validation. The package is designed for simulation, evaluation, prompt conditioning, subgroup analysis, service design, localization evaluation, and benchmark development for AI systems operating in Spain. The release features a layered architecture including a synthetic population layer (persona_core, household_core for stable adult persona structure, household links, economic and territorial context), an LLM interaction layer (persona_views, actor_state_init for compact prompt views and mutable simulation-state scaffolds), and an evaluation layer (benchmark_tasks with replayable tasks, held-out splits, validation metrics). It contains 1,000,000 adult personas, 536,741 households, 6,350,524 LLM-facing views, 1,000,000 actor-state rows, and 1,800 benchmark tasks, totaling 8,889,089 rows and 5.584 GB. It is intended for controlled simulation, evaluation, service design, research prototyping, and benchmark development, but not for surveillance, immigration enforcement, individual-level persuasion, political microtargeting, voter manipulation, eligibility decisions, credit, employment, housing decisions, or claims about real citizens. The dataset includes detailed statistics, evaluation metrics (e.g., regional and age calibration errors), efficiency data (e.g., view-layer token usage), and emphasizes responsible use, ensuring synthetic personas are not treated as real individuals.



