build-small-hackathon/pakistan-notice-helper-traces
收藏资源简介:
该数据集名为NoticeCheck Privacy-Safe Traces,主要用于文本分类任务,专注于安全和诈骗检测,支持英语和乌尔都语。数据集包含NoticeCheck消息审查请求的紧凑、确定性元数据,不涉及隐藏模型推理或自主代理轨迹。数据字段包括:随机追踪ID和时间戳、经过脱敏处理的文本输入或固定模板的图像描述、确定性类别(如快递、银行或FBR)、紧迫性布尔信号、映射模式的确定性摘要、诈骗手法列表以及风险标签和回复草稿策略等扁平化结果列。所有数据单元均为标量字符串、数字或布尔值,无嵌套对象,确保数据易于读取。数据集严格保护隐私,不存储原始消息文本、图像、个人身份信息、模型解释等内容,并通过正则表达式进行脱敏处理,但用户仍应避免提交敏感内容。数据来源包括默认配置的运行时轨迹和示例配置的公开案例,但需注意示例数据不应用于评估类平衡或准确性。数据集存在一定限制,如信号检测近似、频率不代表实际流行度等。
This dataset, named NoticeCheck Privacy-Safe Traces, is designed for text classification tasks, focusing on safety and scam detection, and supports English and Urdu languages. It contains compact, deterministic metadata about NoticeCheck message-review requests, without hidden model reasoning or autonomous-agent trajectories. Fields include: a random trace ID and UTC timestamp, aggressively redacted text input or a fixed-template image description, deterministic categories (e.g., courier, bank, or FBR), a boolean urgency signal, a deterministic summary of the mapped pattern, a list of scam tactics, and flat result columns such as risk label and reply draft policy. All dataset cells are scalar strings, numbers, or booleans, with no dictionaries or nested objects to keep the data easily readable. Privacy is strictly protected by not storing raw message text, images, personal identifiers, model explanations, etc., and using regex-based redaction, though users should opt out for sensitive content. The dataset includes runtime traces in the default configuration and example cases in a separate configuration, but the examples should not be used for estimating class balance or accuracy. Limitations include approximate signal detection and frequencies not reflecting real-world prevalence.




