AIM-Intelligence/guardian-pii-train-data
收藏资源简介:
Guardian-PII训练数据版本是一个用于训练Starfort专用PII(个人可识别信息)模型的版本化数据集集合。该数据集从上游源数据(AIM-Intelligence/guardian-pii-data)中派生,经过混合管道采样和比例平衡,生成训练就绪的数据分片,并按版本冻结以确保可重现性。当前版本v001包含999,338行数据,涵盖两种记录类型:ner-extract(56.85%)和ner-batch-classify(43.15%),涉及两种语言:英语(62.59%)和韩语(37.41%)。数据来源于9个不同的家族(如ai4privacy_1_5m_ko、syvai_en等),包含168种PII实体标签(如GIVENNAME、DATE、SURNAME等),并包含基础示例和注入示例。数据集专门用于训练模型识别和分类PII实体,同时包含负样本(PII不存在)以降低误检率。
Guardian-PII — Training Data Versions is a collection of versioned training datasets used to train Starforts dedicated PII (Personally Identifiable Information) model. Derived from an upstream source dataset (AIM-Intelligence/guardian-pii-data), it undergoes sampling and ratio-balancing via a mixing pipeline to produce train-ready data shards, frozen per version for reproducibility. The current version v001 consists of 999,338 rows, covering two record types: ner-extract (56.85%) and ner-batch-classify (43.15%), and two languages: English (62.59%) and Korean (37.41%). Data is sourced from 9 families (e.g., ai4privacy_1_5m_ko, syvai_en, etc.) and includes 168 PII entity tags (e.g., GIVENNAME, DATE, SURNAME). It contains both base and injected examples, and is designed for training models to detect and classify PII entities, with negative examples (no PII present) to reduce over-detection.




