Run data for "An empirical study of Confidential Compute for frontier AI evaluations"
收藏资源简介:
Paired CC-on / CC-off measurement data from the Pour Demain technical brief"An empirical study of Confidential Compute for frontier AI evaluations:platform overhead and governed egress on Intel TDX and H200 SXM" (June 2026). Contains per-request wall-time, token counts, and payload sizes as ApacheParquet files, plus per-cell summary JSONs, matrix reports, and analysisoutputs (bootstrap CIs, OLS fits, figures). Five measurement campaigns: - phase3/: 18-run primary matrix (1,700 requests), max_tokens sweep, tokens_in sweep, concurrency sweeps, HarmBench cross-corpus replication, GSM8K thinking-mode extrapolation, egress replicates, attestation cell- phase3_v2/: second-deploy max_tokens sweep + concurrency sweep (n=500)- phase3_v3/: third-deploy routing replication- phase3_pysyft/: PySyft governed-egress mechanism validation (§5), comprehensive two-auditor × two-endpoint run (N=195)- phase3_robustness/: inter-deploy variance (2×off, 2×on, 1×attestation) Hardware: 8×H200 SXM, Intel TDX (Emerald Rapids), Tinfoil Containers.Models: GLM-5.1-FP8 (zai-org/GLM-5.1-FP8), Llama-3.1-70B-FP8.Engine: vLLM 0.20.0 with vllm-lens. Analysis harness: https://github.com/pourdemain/cc-benchboxDeployment images: https://github.com/pourdemain/cc-deep-eval
本数据集源自Pour Demain发布的技术简报《面向前沿人工智能评估的机密计算(Confidential Compute)实证研究:Intel TDX与H200 SXM平台开销与受控出口控制》(2026年6月),包含成对的机密计算开启(CC-on)与关闭(CC-off)测量数据。 数据集以ApacheParquet格式存储单请求墙钟时间、Token计数与负载大小,同时附带单单元格摘要JSON文件、矩阵报告与各类分析结果(bootstrap置信区间、普通最小二乘拟合结果、可视化图表)。本次实验共包含5组测量批次: - phase3/:包含18次运行的主矩阵实验(共1700次请求),涵盖max_tokens参数扫描、tokens_in参数扫描、并发度扫描、HarmBench跨语料库复现、GSM8K思维模式外推、出口控制重复实验、认证单元格相关实验; - phase3_v2/:第二次部署的max_tokens参数扫描与并发度扫描实验(样本量n=500); - phase3_v3/:第三次部署的路由复现实验; - phase3_pysyft/:PySyft受控出口控制机制验证实验(对应第5章),包含双审计员×双端点的全面实验(总样本量N=195); - phase3_robustness/:跨部署方差实验(2次机密计算关闭、2次机密计算开启、1次认证实验); 硬件配置:8块H200 SXM加速卡、Intel TDX(基于Emerald Rapids平台)、Tinfoil Containers。 测试模型:GLM-5.1-FP8(仓库标识:zai-org/GLM-5.1-FP8)、Llama-3.1-70B-FP8。 推理引擎:搭载vllm-lens扩展模块的vLLM 0.20.0版本。 分析工具链开源地址:https://github.com/pourdemain/cc-benchbox 部署镜像开源地址:https://github.com/pourdemain/cc-deep-eval



