NativePort Web-Access API Benchmarks — run 2026-08-05
收藏资源简介:
An immutable, checksummed archival release of one NativePort benchmark run: run 2026-08-05, covering 22 commercial web-access APIs across 13 capabilities, 67 provider x capability scorecards and 297 individual metric measurements. This record exists so that a figure quoted from these benchmarks can be traced, years later, to the exact bytes it was read from. WHAT IS IN THE RECORD nativeport-web-access-api-benchmarks-2026-08-05.zip (36397 bytes, SHA-256 8a52099e3b8b27774cf45e72af38929543ab829b49148a56ce5d5523e8de47e9) — the complete release: four data files, the licence, a citation file and a per-file SHA-256 manifest. nativeport-web-access-api-benchmarks-2026-08-05.zip.sha256 — the archive digest on its own, for sha256sum -c. README.md — schema reference, provenance table and integrity instructions, extracted from the archive so it can be read without downloading it. CITATION.cff — machine-readable citation metadata (Citation File Format 1.2.0). LICENSE.txt — the complete CC BY 4.0 legal code. DATA SHAPES data/benchmarks.csv — 297 tidy rows, 23 columns, one row per provider x capability x metric. data/metric_rows.jsonl — the same 297 rows as JSON Lines, with JSON number and boolean types. data/benchmarks.jsonl — 67 records, one per scorecard, metrics nested as a list of structs. data/summary.json — computed counts, per-capability and per-metric inventories, and the SHA-256 of each data file as generated. PROVENANCE AND FIXITY Every file was generated by a deterministic standard-library program from a single public JSON snapshot, https://nativeport.ai/evals.json (schema version 1, 89894 bytes, SHA-256 f46bf416803d5adf696c66ac14a6cbf06f4dfa39727cca164e8f3872abd6ed9b). Identical input bytes produce byte-identical output; no wall-clock timestamp is written anywhere in the release. Numeric values are carried across as exact source tokens — the generator verifies that each emitted number serialises character-for-character back to the token it read, and aborts rather than emit a rounded stand-in. To verify a download: sha256sum -c nativeport-web-access-api-benchmarks-2026-08-05.zip.sha256, then unzip and run sha256sum -c CHECKSUMS.sha256 from the extracted nativeport-web-access-api-benchmarks-2026-08-05/ directory. WHO PRODUCED THIS, AND WHAT IT IS NOT NativePort produced these measurements from its own first-party benchmark runs. NativePort operates a commercial gateway that routes to many of the providers scored here, so this is not an independent third-party evaluation and must not be cited as one. Weak and last-place results are archived unchanged: three scorecards in this run sit below 2.0 out of 10, and two record a 100% error rate. No provider named in this release has endorsed, reviewed or sponsored it. Each capability is scored on a versioned task corpus held fixed across every provider, and four dimensions are recorded per provider x capability pair: a capability-specific quality metric, median wall-clock latency measured at the runner, track-specific cost computed from the provider's real metered price, and error rate across the run. Metrics that need judgment rather than string comparison are graded by an LLM panel against written rubrics. The measurement protocol, including what NativePort declines to claim, is documented at how we measure; the ranked human-readable tables are at https://nativeport.ai/leaderboards/. SCOPE AND LIMITATIONS OF THIS RELEASE Point in time. Every value was measured on 2026-08-05. Provider behaviour, pricing and anti-bot posture change; this record is a fixed observation, not a current statement. Partial coverage. 22 of the 32 providers in the source catalog carry scorecards in this run. Absence means not measured in this run — not failed, and not unavailable. Composites are capability-local. Averaging a provider's composite scores across capabilities produces a number with no defined meaning. Thin capabilities. act_agent and watch have one scored provider each; read rank together with rank_of. Cross-capability latency and cost comparisons are invalid, and the three cost denominators (per call, per useful result, per successful page) must not be mixed in one ordering. Judged metrics carry model bias, being LLM-graded against rubrics rather than human-annotated. Not measured at all: uptime, SLA conformance, throughput ceilings, regional performance, concurrency behaviour and long-run stability. Cost figures are measurements taken from metered prices at run time, not a quote, an offer or a rate card. LICENCE Licensed by NativePort under the Creative Commons Attribution 4.0 International licence, SPDX CC-BY-4.0 (https://creativecommons.org/licenses/by/4.0/). The full legal code ships as LICENSE.txt. NativePort licenses the compilation — its selection, arrangement, schema and documentation — together with the measurements NativePort itself produced and any database rights NativePort holds, and only to the extent NativePort holds them. CC BY 4.0 grants no patent or trademark rights: provider, product and company names appearing in the data are the marks of their respective owners, used for identification and factual comparison only, and are not licensed here. The material is offered as-is, without warranties or conditions of any kind. RELATED RECORDS The same benchmark run is distributed in Hugging Face dataset form at https://huggingface.co/datasets/nativeport/web-access-api-benchmarks. That copy tracks the Hub's dataset-card and loader conventions; this record is the fixed, checksummed archival form.



