Twin-2K-500
收藏资源简介:
Twin-2K-500是一个大规模公开数据集,旨在通过超过500个问题的回答,捕捉个人行为的丰富和全面视图。数据集由美国2,058名代表性参与者回答,涵盖人口统计、心理、经济、个性、认知等方面的测量,以及行为经济学实验的复制品和定价调查。数据集通过四次调查收集,每次调查约2.42小时。通过公开提供完整数据集,旨在为基于LLM的人物模拟的开发和基准测试建立一个有价值的测试平台。此外,数据集的独特广度和规模也使其能够进行广泛的社会科学研究,包括跨结构相关性和异质处理效应的研究。
Twin-2K-500 is a large-scale public dataset designed to capture a rich and comprehensive portrait of individual behaviors through responses to over 500 questions. The dataset is answered by 2,058 representative participants in the United States, covering measurements across demographic, psychological, economic, personality, and cognitive domains, as well as replications of behavioral economics experiments and pricing surveys. The dataset was collected across four survey waves, each lasting approximately 2.42 hours. By publicly releasing the full dataset, this work aims to establish a valuable testbed for the development and benchmarking of LLM-powered human simulation. Furthermore, the unique breadth and scale of the dataset also enable a wide range of social science research, including studies on cross-structure correlations and heterogeneous treatment effects.

- 1Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions哥伦比亚大学 · 2025年



