vLLM-inference
收藏官方服务:
资源简介:
A high-throughput and memory-efficient inference and serving engine for LLMs
面向大语言模型(Large Language Model,LLM)的高吞吐量且内存高效的推理与服务引擎
创建时间:
2025-03-01

A high-throughput and memory-efficient inference and serving engine for LLMs
面向大语言模型(Large Language Model,LLM)的高吞吐量且内存高效的推理与服务引擎