遇见数据集

MultiPII: A Multimodal Benchmark for Personal Identifiers Detection in Real-World Documents

收藏
Zenodo2025-08-19 更新2026-05-26 收录
官方服务:

资源简介:

MultiPII is a curated image–text (multimodal) dataset for detecting and redacting personally identifiable information (PII) in real-world documents and social posts. The corpus contains 197 images (scans, screenshots, document photos, and posts) with paired annotations of PII. Each item includes (i) the source file/URL and (ii) a list of PII annotations recorded as Type: Value (e.g., Passport Number: [REDACTED], Date of Birth: [REDACTED], Name: [REDACTED]). The dataset is designed for benchmarking multimodal PII extraction, privacy-aligned redaction, and error analysis (missed identifiers, over-redaction, hallucinated PII). We release MultiPII to support privacy-preserving NLP/vision research, evaluation of LLM/VLM pipelines (OCR→text PII detection + visual cues), and the development of redaction-aware metrics.

提供机构:
Zenodo
创建时间:
2025-08-19
二维码
社区交流群
二维码
科研交流群
商业服务