PANORAMA
收藏资源简介:
PANORAMA是一个大规模的合成语料库,包含384,789个样本,源自9,674个合成个人资料,旨在模拟网络环境中自然出现的个人身份信息(PII)和敏感数据的分布、多样性和上下文。数据集涵盖了多种内容类型,包括维基风格的文章、社交媒体帖子、论坛讨论、在线评论、评论和市场列表等。数据集和代码公开发布,为隐私风险评估、模型审计和隐私保护的大型语言模型(LLMs)的开发提供了必要的资源。
PANORAMA is a large-scale synthetic corpus comprising 384,789 samples derived from 9,674 synthetic user profiles. It is designed to mimic the distribution, diversity and contextual characteristics of personally identifiable information (PII) and sensitive data that naturally emerge in online environments. The corpus covers diverse content types, including Wikipedia-style articles, social media posts, forum discussions, online reviews, comments, and marketplace listings, among others. The dataset and its accompanying code are publicly released, providing essential resources for privacy risk assessment, model auditing, and the development of large language models (LLMs) for privacy protection.

- 1PANORAMA: A synthetic PII-laced dataset for studying sensitive data memorization in LLMs微软 · 2025年



