DeepSeek 1M Context Benchmark: Retrieval Accuracy, Latency, and Cost — Dataset v1.0.0
收藏资源简介:
This checksum-bound dataset contains real DeepSeek API measurements over deterministic synthetic English record corpora. The run was conducted from 2026-08-06T20:17:02.706Z to 2026-08-07T00:07:44.737Z under frozen protocol deepseek-v4-long-context-retrieval-v1.1.0. The release contains 344 sanitized terminal records: 288 U.S. client-vantage primary accuracy cases, 20 excluded pilot calls, and 36 matched India client-vantage validation calls. It compares deepseek-v4-flash and deepseek-v4-pro across four retrieval and synthesis task families, beginning/middle/end target positions, three deterministic repeats, and provider-counted prompt tiers of 32K, 128K, 512K, and approximately 950K tokens. In the primary 288-case denominator, Flash achieved 71/144 strict exact matches (49.3%) and Pro achieved 81/144 (56.3%). U.S. primary end-to-end latency was 12.748 seconds at p50 and 162.295 seconds at p95. The dated all-role cache-miss cost upper bound was USD 39.913229 with complete usage coverage. Only the 288 U.S. primary rows enter the published accuracy denominator; the 20 pilot rows and 36 India client-vantage rows are excluded from it. AWS regions identify client network vantages, not user populations or provider-hosting locations. Latency combines client-network and service effects from one narrow run window and is not a reliability study. The benchmark uses synthetic English records and strict JSON grading, so it does not measure general model quality. Each case received one attempt with no automatic retries. Pricing is frozen to August 6, 2026. The 256-token output cap was selected for this study, not claimed as the model maximum. Model routing and fingerprints are dated observations. “1M” refers to the documented model context capacity; the largest tested tier was approximately 950K provider-counted prompt tokens, not one million input tokens. The deposit includes sanitized CSV and JSONL measurements, frozen methodology and protocol files, derived analysis tables, original charts and chart-source data, data definitions, QA evidence, and integrity metadata. Raw prompts, raw provider responses, tokenizer artifacts, credentials, cloud identifiers, private paths, and account data are excluded. The sanitized data, derived tables, original charts, chart-source data, and dataset documentation are licensed under CC BY 4.0. Reproducibility software is excluded from this deposit and remains separately available under MIT in the linked GitHub release.



