MultiPII: A Multimodal Benchmark for Personal Identifiers Detection in Real-World Documents
收藏资源简介:
MultiPII is a curated image–text (multimodal) dataset for detecting and redacting personally identifiable information (PII) in real-world documents and social posts. The corpus contains 197 images (scans, screenshots, document photos, and posts) with paired annotations of PII. Each item includes (i) the source file/URL and (ii) a list of PII annotations recorded as Type: Value (e.g., Passport Number: [REDACTED], Date of Birth: [REDACTED], Name: [REDACTED]). The dataset is designed for benchmarking multimodal PII extraction, privacy-aligned redaction, and error analysis (missed identifiers, over-redaction, hallucinated PII). We release MultiPII to support privacy-preserving NLP/vision research, evaluation of LLM/VLM pipelines (OCR→text PII detection + visual cues), and the development of redaction-aware metrics.



