遇见数据集

Ollama vs llama.cpp vs LM Studio: Controlled Runtime Comparison on One RTX 5080

收藏
Zenodo2026-08-30 更新2026-10-01 收录
官方服务:

资源简介:

Decode, prefill, VRAM, RAM working set, time-to-first-token and load time for Ollama, llama.cpp and LM Studio measured on a single retail NVIDIA RTX 5080. Experimental control. Most published runtime comparisons do not hold the model file constant, so the reported delta may reflect an uncontrolled quantization difference rather than the runtime. Here the same GGUF files are shared between applications via hard links and verified by sha256, so each runtime provably decodes identical bytes. Method. Temperature 0, seed 42, and a per-iteration prompt nonce; without the nonce, warm repetitions hit the prompt cache and report approximately 60,000 tok/s of "prefill" that measures nothing. Fresh server per model, wall-clock cross-check on every run, llama-bench as a third instrument on the 3B model. Results. 2026-08-20: llama.cpp b10507 decodes 2-6% faster than Ollama 0.32.1 on dense models. 2026-08-26: Ollama 0.32.15 decodes 5-11% faster than LM Studio 0.4.21 (351 vs 316 tok/s on Llama 3.2 3B), while LM Studio prefills faster. Ollama gained 3-10% decode against itself between 0.32.1 and 0.32.15 following a vendored engine update, a delta larger than several cross-runtime gaps. Stated caveat. Ollama's gpt-oss blob carries its own "gptoss" architecture tag, so llama.cpp cannot be pointed at identical bytes for that model. Using the upstream GGUF it decodes 14% faster, but this is a conversion difference and is labelled as such rather than reported as a runtime result. Negative result retained. A background indexing process was found to depress 3B decode by approximately 15% mid-capture; it was suspended and all published cells re-measured in the quiet window. Canonical page, full method and change log: https://techfuelhq.com/data/llm-server-compare/

提供机构:
Zenodo
创建时间:
2026-08-30
二维码
社区交流群
二维码
科研交流群
商业服务