sie-task-evidence
收藏资源简介:
SIE task evidence 是一个用于重现 superlinked.com 上任务页面结果的数据集。它记录了每个任务的实际输入和模型响应,按任务组织为独立文件夹。每个文件夹包含三个部分:inputs/(精确的文档、查询、图像或音频输入)、calls.json(每次调用的请求、响应、状态和计时详情)以及 manifest.json(端点、模型 ID、服务修订版和运行日期)。特别地,redact 子文件夹用于个人数据检测基准测试,包含 660 篇英文文档(来自 gretelai/synthetic_pii_finance_multilingual 数据集的测试集),每篇文档标注了黄金标准 PII 跨度(共 1792 个跨度),并记录了五个模型臂(SIE、Microsoft Presidio、OpenAI Privacy Filter、GPT-6 Luna、Claude Haiku 4.5)的响应、结果、令牌计数和成本。数据集的目的是让任何人都能使用 Python 标准库从头重现已发布的评估数字,而无需 API 密钥或推理支出。数据集仅保证这些数字来自真实运行,并不代表输入对用户数据的适用性。输入来源公开,许可证因来源而异(redact 部分采用 Apache 2.0)。
SIE task evidence is a dataset for reproducing results from task pages on superlinked.com. It records actual inputs and model responses for each task, organized into independent folders per task. Each folder contains three parts: inputs/ (exact documents, queries, images, or audio inputs), calls.json (request, response, status, and timing details for each call), and manifest.json (endpoint, model ID, service revision, and run date). Notably, the redact subfolder is used for personal data detection benchmarking, containing 660 English documents (from the test set of the gretelai/synthetic_pii_finance_multilingual dataset), each annotated with gold-standard PII spans (1792 spans in total), and records responses, results, token counts, and costs for five model arms (SIE, Microsoft Presidio, OpenAI Privacy Filter, GPT-6 Luna, Claude Haiku 4.5). The dataset aims to allow anyone to reproduce published evaluation numbers from scratch using Python standard library, without requiring API keys or inference expenses. The dataset only guarantees these numbers come from real runs and does not represent the applicability of inputs to user data. Input sources are public, and licenses vary by source (the redact part is under Apache 2.0).
SIE task evidence 数据集概述
基本信息
- 数据集名称:SIE task evidence
- 许可证:other
- 标签:evaluation、reproducibility
- 数据集地址:https://huggingface.co/datasets/superlinked/sie-task-evidence
数据集用途
该数据集记录了 superlinked.com 任务页面背后的输入和模型响应,每个任务对应一个文件夹。任务页面上发布的每个数字均由针对 https://api.superlinked.com 的真实记录运行产生;对于某些任务(如 typed-decisions),则是针对自托管 SIE 服务器在 GPU 上固定提交(pinned commit)运行产生,该提交记录在任务的 manifest.json 中。
该数据集的核心目的是让任何人都能无需 API 密钥、无需消耗任何推理即可重新推导出已发布的数字。
使用方法
每个任务的可运行示例位于公开仓库 superlinked/sie 的 examples/ 目录下。每个示例包含运行器(runner)、评分器(scorer)和 README。操作流程如下:
git clone https://github.com/superlinked/sie cd sie/examples/<task> python3 fetch.py # 下载此数据集,版本为该示例所固定的修订版 python3 score.py # 复现已发布页面上的数字
特点:
- 仅使用标准库,无依赖安装、无 Hugging Face token、无 API 密钥、无推理开销
- 每个示例固定一个提交 SHA 而非
main,确保所评分的字节与编写时一致 fetch.py携带manifest.json的摘要,进而固定calls.json和每个输入,因此缺失或被篡改的文件会导致获取失败,而非静默产生不同数字
目录结构
<task>/ inputs/ 使用的确切文档、查询、图像或音频 calls.json 每次调用一条记录:请求、响应、状态、计时 manifest.json 端点、模型 id、服务修订版、运行日期
数据集的证明范围
- 可以证明:任务页面上的数字来自实际发生过的运行,且任何人都能从相同的字节重新计算
- 不能证明:这些输入能代表你的数据;相同模型在你的文档上会得到相同分数;未来模型修订版行为一致
- 若页面展示的是更大记录集的选取部分,页面会注明,完整集在此处提供
/redact 子集
redact/ 保存了 superlinked.com/redact 背后的运行记录(来源:SOURCES.md),记录日期为 2026-09-30。
- 测量指标:覆盖率召回率(coverage recall)——被某方案掩码的范围内个人数据跨度(HIPAA Safe Harbor 标识符加凭据)中每个非空格字符被掩码的比例
- 规模:660 篇文档,1,792 个范围内跨度,五个方案
- 五个方案:SIE(
urchade/gliner_multi_pii-v1和numind/NuNER_Zero在https://api.superlinked.com上组合)、Microsoft Presidio、OpenAI Privacy Filter、GPT-6 Luna、Claude Haiku 4.5
redact 目录结构
redact/ manifest.json 运行日期、端点、模型、标签、提示词、方案设置、每个文件的 sha256 inputs/ gretel-main.jsonl:660 篇文档及其黄金跨度 rows/ 每篇文档每方案一行 JSONL,该方案的记录响应 results/ 研究图表、token 计数、吞吐量及每方案价格
examples/redact 在固定修订版获取此文件夹,并使用 Python 标准库重新推导每个已发布图表。
redact 数据来源
- 文档为
gretelai/synthetic_pii_finance_multilingual英文测试集的 660 行,修订版7b844d16738527a04264f50214cb426a4cea0897,由 Gretel.ai 提供,采用 Apache License 2.0 许可 - 文本和黄金跨度未经修改地重新分发;字段已重命名(
generated_text→text,document_type→domain,pii_spans→gold,每个跨度保留start、end和label) - 其中的个人数据为合成数据
来源与许可
输入来自公共来源,每个任务的 manifest.json 记录每个输入的来源。许可证因来源而异,目前正在审查中。若你是此处某内容的权利持有人并希望移除,请在本数据集上发起讨论。





