An open-source inference engine that introduced PagedAttention for efficient KV cache memory management, becoming a standard choice for serving LLMs at scale.
An open-source inference engine using PagedAttention for efficient KV cache memory management.
An open-source inference engine that introduced PagedAttention for efficient KV cache memory management, becoming a standard choice for serving LLMs at scale.