LLM Latency Tracker: measured latency and uptime for AI inference APIs (2026-07-23 to 2026-08-15)
收藏资源简介:
Independent, continuously measured latency and availability for AI inference API providers, aggregated by day. Covers 45 providers across 4 regions (ap-tokyo, eu-hetzner, sa-east, us-central), built from 1,210,851 raw probes collected between 2026-07-23 and 2026-08-15. Probes run every five minutes from separate network locations and are never routed through a gateway or an aggregator, so the numbers describe the providers themselves rather than a proxy in front of them. Network probes measure DNS → TCP → TLS → time to first byte; inference probes measure time to first token on a real completion request. Files. daily_aggregates.csv — one row per date, provider, region and probe type, with p50/p95, sample count and success rate. rankings.json — the machine-readable snapshot published live at llmlatency.dev. Percentiles are nearest-rank, identical to the ones shown on the site. Limitations. Vantage points are cloud data centres, not consumer networks, so absolute values are lower than an end user would see; the comparison between providers is the meaningful part. Provider coverage changes over time as APIs appear and shut down. Live data, full methodology and the raw-data retention policy: llmlatency.dev.



