Skip to main content

What is vLLM?

vLLM is an open-source inference engine for large language models (LLMs). Organizations use vLLM for higher throughput and greater memory efficiency of their LLMs, supporting scalable machine learning (ML) operations. vLLM is designed to overcome the limitations of standard LLM memory management in the GPU’s key-value cache (KV cache) by using the PagedAttention algorithm. vLLM works with a wide range of existing models, hardware accelerators, and parallelization techniques.

Why does LLM inference need a dedicated serving engine?

vLLM architecture on Amazon EKS

Implementing large language models within the business means that you have to consider inference latency, or request-reply cycle times, for a growing user base. While LLMs can deliver productivity gains, these can be minimized when users wait long times for replies. LLMs must be able to handle a growing number of concurrent requests in acceptable timeframes.

Conventionally, LLMs can struggle with time-to-inference due to inefficient memory allocation. Each sequence, or user request and associated token stream, is allocated to a predefined, contiguous block in the GPU’s KV cache, up to the maximum sequence length. The contents of the key-value cache provide context for the LLM to generate subsequent responses. The all-at-once assignment of sequences to cache blocks causes inefficiency, as most sequences fall far short of the maximum length. Some blocks remain unused during inference, and the memory becomes fragmented. This leads to performance degradation, affecting applications that require concurrent processing, such as chatbots.

The dedicated vLLM engine overcomes the memory management challenges of traditional KV cache through virtualization and improves inference latency. The vLLM inference engine offers batch processing, optimized memory usage, and high-throughput inference.

What is PagedAttention?

PagedAttention is a memory management technique that maps data in contiguous logical blocks to non-contiguous physical blocks. This is the key differentiator between vLLM and other inference engines. The PagedAttention technique solves the problem of inefficient use of the key value cache, whereby the conventional inference engine stores key-value tensors sequentially in preallocated slots. The PagedAttention memory virtualization technique is similar to the operating system’s virtual memory and paging mechanisms.

PagedAttention breaks sequences of data into multiple physical blocks. Consider these physical key-value blocks.

PagedAttention physical key-value blocks table

Instead of holding the entire sequence in one block, it is split across several non-contiguous blocks. Here, the only wasted memory space is in the very last block in the sequence, rather than wasting memory at the end of every sequence. Consequently, teams can deploy the memory-efficient inference engine in production systems with significantly lower memory waste.

How does vLLM work?

vLLM enables LLMs to infer more efficiently through PagedAttention and a coordinated, distributed architecture. It supports popular models, including Llama, Qwen, and DeepSeek. By using the PagedAttention algorithm to manage the KV cache, vLLM allows ML teams to run models with fewer GPUs. Here are some key features of vLLM.

Continuous batching

Continuous batching is the vLLM’s ability to process new queries as soon as a decode slot becomes available, rather than waiting for all previous queries to complete. vLLM adds new requests to an existing request batch to maximize GPU utilization with variable request lengths. These new requests can be added during generation.

Prefill and decode phases

LLM inference includes the prefill and decode phases. In the prefill phase, the LLM processes tokens to determine the key-value tensors to store in the KV cache. The decode phase generates tokens iteratively, using the KV cache. The decode phase is complete when the sequence ends or the maximum number of tokens is met.

Distributed inference

Distributed inference is the process of distributing the workload of a generative AI model to multiple GPUs. Distributed inference is necessary when a model doesn’t fit on a single GPU and is preferred when the model has high inference latency, to fit the model and speed up inference. While vLLM has a centralized scheduler, KV cache manager, CPU block allocator, and GPU block allocator, the LLM’s work is distributed across multiple GPUs.

vLLM supports various methods of distributed inferencing:

  • Pipeline parallelism assigns different model layers to different GPUs. Each GPU passes the results on to the next after it has finished processing. This method helps improve throughput and is useful when a model is too big to fit on a GPU.
  • Tensor parallelism distributes the model weight matrices of each layer across multiple GPUs, effectively splitting layers. vLLM uses tensor parallelism when a tensor layer is too large to fit in a single GPU’s memory.
  • Expert parallelism allows you to distribute a mixture of experts (MoE) architecture across different GPUs. MoE is a machine learning method that breaks a neural network into multiple smaller expert sub-networks for distribution across GPUs. Using expert parallelism, vLLM can process incoming requests and route tokens to the most well-suited expert in the MoE network.

Quantization

Quantization compresses the tensor weights into a lower-precision format to reduce memory space. vLLM supports several quantized formats, including INT8, INT4, FP8, GPTQ, and AWQ. This allows the inference engine to accelerate prefill and decoding while maintaining minimal accuracy loss.

What are the key capabilities of vLLM?

vLLM enables teams to deploy, operate, and scale LLMs on GPUs and reduce hardware costs.

OpenAI-compatible APIs

ML teams can serve models directly through the OpenAI-compatible HTTP server. This allows vLLM to act as a drop-in replacement for OpenAI’s backend. By installing vLLM and starting the server, you can access OpenAI APIs such as the Completions API, Responses API, and Chat Completions API from an HTTP client.

Prefix caching

Prefix caching allows vLLM to reuse prefixes stored in the KV cache when it encounters the same leading tokens or phrases in requests. vLLM automates the process so the model doesn’t need to recalculate the entire request. This is helpful when processing long document queries or multiround conversations in chatbots. For example, vLLM allows the model to reuse chat history in the KV cache when generating a response.

Multi-LoRA serving

vLLM supports Multi-LoRA in deployment. Multi-LoRA is a variant of the Low-Rank Adaptation (LoRA) technique for fine-tuning large models. Instead of retraining the entire model, LoRA freezes the base weights and injects a trainable adapter into the model. Multi-LoRA works similarly but manages multiple adapters at the same time. This allows multiple fine-tuned adapters to share one GPU, with the adapter selected per request based on inference requirements. Consequently, teams can serve many custom model variants on a single GPU.

What are the use cases for vLLM?

vLLM helps organizations overcome latency issues when scaling LLMs.

Real-time inference APIs

vLLM provides real-time inference APIs that allow more efficient use of the underlying hardware. Organizations can deploy models enabling high-concurrency chat, code generation, or summarization that fully utilize the GPU even under varying request loads. These workloads leverage techniques such as continuous batching, PagedAttention, and quantization to accelerate inference.

Multi-tenant and multi-variant serving

Deploying multi-tenant AI applications, servicing multiple users or applications, conventionally requires serving an LLM from a single GPU. This often results in latency issues when scaling, which can be overcome with vLLM and parallelism techniques. vLLM also enables organizations to host multiple model variants on a single GPU using techniques such as Multi-LoRA.

Offline batch inference

vLLM supports offline or asynchronous inference, allowing datasets to be processed in batches. For example, vLLM is suitable for document processing, dataset generation, and evaluation pipelines. These workloads prioritize high throughput instead of low latency when generating predictions for non-immediate analysis.

What hardware does vLLM run on?

vLLM on Amazon EC2 Trainium

vLLM supports CPUs, GPUs, and AI accelerators from various manufacturers.

For example, vLLM supports:

  • NVIDIA CUDA GPUs, allowing you to run CUDA kernels and supporting FlashAttention.
  • AMD ROCm GPUs, allowing you to run inference with the ROCm software stack.
  • AWS Inferentia2 and Trainium AWS-designed AI accelerators, allowing you to integrate with the AWS Neuron full SDK, and supporting continuous batching.

Other hardware that vLLM supports includes Intel CPUs/GPUs, ARM, Google TPU, IBM Spyre, and Huawei Ascend NPU.

When serving a model on vLLM with AWS Neuron, be mindful of the first-load latency as Neuron compiles the model into an artifact before inference. To reduce latency, implement pre-compilation and cache the artifacts.

What are the key deployment considerations for vLLM?

While vLLM provides high-throughput inference at scale, you should consider several factors to help with a secure and efficient deployment.

Instance sizing

Plan the GPU requirement to fit the workload footprint. This includes the model weights, KV cache, and activation memory typical of an LLM. If you’re deploying a large model, such as Meta’s Llama 3.3 70B, consider applying optimization techniques such as tensor or pipeline parallelism. Additionally, you can downsize the model with quantization to fit the GPU and meet the sizing requirements.

Latency vs. throughput trade-offs

Conventionally, continuous batching improves overall throughput but increases time-to-first-token (TTFT) for high-concurrency workloads. This is because the inference engine attempts to process the prefill and decode stages within the same instance, with prefill blocking the decode steps, increasing latency. vLLM allows ML engineers to implement disaggregated prefill, which places prefill and decode at separate instances. Additionally, you can decrease the max-num-seqs in the model engine to reduce KV space usage, with the trade-off being reduced throughput.

Authentication and network security

vLLM is designed without default security measures. Without built-in authentication, third parties can access the models through the exposed API. When deploying vLLM for enterprise, restrict public network access to the inference engine. Use protective measures, such as access controls, on media, containers, and endpoints that vLLM connects to. For example, you can place vLLM behind Amazon API Gateway to enable robust authentication for non-streaming workloads. For streaming workloads, you can use Application Load Balancer.

How can AWS support your vLLM deployment requirements?

AWS supports vLLM-based inference workloads, with a range of infrastructure and managed services. Consider these services:

  • Amazon Bedrock is the platform for building generative AI applications and agents at a production scale. Amazon Bedrock supports vLLM inference, including Multi-Low-Rank Adaptation (Multi-LoRA) serving of popular open-source models.
  • Amazon EC2 Inf2 instances are purpose built for deep learning inference. They deliver high performance at the lowest cost in Amazon EC2 for generative AI models, including LLMs and vision transformers. To run vLLM on Inf2 you use the AWS Neuron SDK.
  • Amazon SageMaker AI is a fully managed service that brings together the most comprehensive set of AI tools and capabilities to enable high-performance, low-cost AI model development for any use case. Similar to Bedrock, Amazon SageMaker AI supports vLLM inference.

Get started with vLLM inference on AWS by creating a free account today.

Browse all cloud computing concepts

Browse all cloud computing concepts content here:

Loading
Loading
Loading
Loading
Loading

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages