遇见数据集

Run data for "An empirical study of Confidential Compute for frontier AI evaluations"

收藏
Zenodo2026-06-01 更新2026-06-05 收录
官方服务:

资源简介:

Paired CC-on / CC-off measurement data from the Pour Demain technical brief"An empirical study of Confidential Compute for frontier AI evaluations:platform overhead and governed egress on Intel TDX and H200 SXM" (June 2026). Contains per-request wall-time, token counts, and payload sizes as ApacheParquet files, plus per-cell summary JSONs, matrix reports, and analysisoutputs (bootstrap CIs, OLS fits, figures). Five measurement campaigns: - phase3/: 18-run primary matrix (1,700 requests), max_tokens sweep, tokens_in sweep, concurrency sweeps, HarmBench cross-corpus replication, GSM8K thinking-mode extrapolation, egress replicates, attestation cell- phase3_v2/: second-deploy max_tokens sweep + concurrency sweep (n=500)- phase3_v3/: third-deploy routing replication- phase3_pysyft/: PySyft governed-egress mechanism validation (§5), comprehensive two-auditor × two-endpoint run (N=195)- phase3_robustness/: inter-deploy variance (2×off, 2×on, 1×attestation) Hardware: 8×H200 SXM, Intel TDX (Emerald Rapids), Tinfoil Containers.Models: GLM-5.1-FP8 (zai-org/GLM-5.1-FP8), Llama-3.1-70B-FP8.Engine: vLLM 0.20.0 with vllm-lens. Analysis harness: https://github.com/pourdemain/cc-benchboxDeployment images: https://github.com/pourdemain/cc-deep-eval

提供机构:
Zenodo
创建时间:
2026-06-01
二维码
社区交流群
二维码
科研交流群
商业服务