Xinference is a fast-growing open-source inference engine in the AI ecosystem -- 9,300+ GitHub stars, 140+ contributors, and 8.3 M downloads. The Xinference enterprise platform builds on this foundation to deliver production-ready LLM serving on AWS GPU infrastructure.
Key Features
Unified OpenAI-compatible API -- REST and gRPC; swap commercial LLMs for any open-source model with a one-line config change.
Multi-backend optimisation -- vLLM, SGLang, llama.cpp, and TensorRT-LLM backends with automatic selection per workload.
Continuous batching & prefix caching -- high-throughput serving with automatic optimisation for shared system prompts and RAG context.
Multi-modal model support -- LLMs, embedding, reranking, image generation, and speech models under a single platform.
Distributed deployment -- scale across multiple GPUs and nodes with built-in cluster management.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by GPU-hours, so cost accumulates based on how long you run models on each GPU type. Each dimension maps to a specific NVIDIA GPU class and memory size: T4 16GB, L4 24GB, A10G 24GB, A100 80GB, H100 80GB, H200 141GB, and B200 180GB. The larger multi-GPU dimensions bill for eight-GPU configurations. You choose the GPU that fits your workload, and your total scales with the hours you use across the selected types. An additional N/A unit also appears in the table.
Top-of-mind questions for buyers
What does one GPU-hour unit measure, and how is it counted?
One GPU-hour is one hour a GPU runs your models. Cost accrues per hour of active runtime on the GPU class you select. The multi-GPU dimensions bill eight-GPU configurations, so each hour reflects the full eight-card node running together, not a single card.
Am I charged when GPUs sit idle or when models are not running?
Charges accumulate based on GPU-hours of runtime. The platform auto-scales replicas to track traffic, scaling out during spikes and back down afterward. Running fewer replicas reduces accrued hours. Underlying AWS infrastructure fees may still apply separately when instances remain provisioned.
If I run several GPU types at once, how do the charges combine?
Each GPU dimension bills independently by its own accumulated hours. Your invoice adds the hours from every GPU class you run. Larger multi-GPU nodes drive more cost per hour because they meter eight-GPU configurations. You choose the mix that fits your workload.
xinference.co
Helpful?
Vendor refund policy
NA
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
HAMi is an open-source Kubernetes middleware providing GPU compute and memory isolation, flexible GPU slicing, and topology-aware scheduling maximizing utilization for AI inference and training workloads across NVIDIA GPUs and AWS Neuron devices; part of the CNCF ecosystem. Works with NVIDIA GPU Operator, vLLM Production Stack, and Xinference.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.