Xinference is a fast-growing open-source inference engine in the AI ecosystem -- 9,300+ GitHub stars, 140+ contributors, and 8.3 M downloads. The Xinference enterprise platform builds on this foundation to deliver production-ready LLM serving on AWS GPU infrastructure.
Key Features
Unified OpenAI-compatible API -- REST and gRPC; swap commercial LLMs for any open-source model with a one-line config change.
Multi-backend optimisation -- vLLM, SGLang, llama.cpp, and TensorRT-LLM backends with automatic selection per workload.
Continuous batching & prefix caching -- high-throughput serving with automatic optimisation for shared system prompts and RAG context.
Multi-modal model support -- LLMs, embedding, reranking, image generation, and speech models under a single platform.
Distributed deployment -- scale across multiple GPUs and nodes with built-in cluster management.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay by GPU-hour, so billing tracks how long you run instances rather than tokens or requests. Each dimension maps to a specific GPU type: T4 16GB, L4 24GB, A10G 24GB, and multi-GPU configurations for A100 80GB, H100 80GB, H200 141GB, and B200 180GB. You accumulate hours only on the GPU types you actually use. Costs scale with runtime and with the class of hardware you select. Mixing GPU types is possible, and each type bills independently under its own dimension.
Top-of-mind questions for buyers
What does one GPU-hour cover, and how do the multi-GPU dimensions differ from the single-GPU ones?
One GPU-hour is one hour of runtime for the listed GPU type. T4, L4, and A10G dimensions bill single-GPU instances. The A100, H100, H200, and B200 dimensions each bill an eight-GPU configuration, so one accumulated hour covers all eight GPUs running together.
Am I charged when my GPU instances are idle or powered off?
You accumulate hours only while instances run. Billing tracks runtime per GPU type, so stopped instances stop adding hours under that dimension. Underlying AWS storage or reservation fees may still apply separately, but the GPU-hour meter counts running time only.
If I run several GPU types at once, how do the charges combine on my bill?
Each GPU type bills independently under its own dimension. Hours on T4 do not affect hours on H100. Your total is the sum of hours accumulated across every GPU type you run. The hardware class you select and total runtime together drive cost.
xinference.co+1
Helpful?
Vendor refund policy
NA
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
HAMi is an open-source Kubernetes middleware providing GPU compute and memory isolation, flexible GPU slicing, and topology-aware scheduling maximizing utilization for AI inference and training workloads across NVIDIA GPUs and AWS Neuron devices; part of the CNCF ecosystem. Works with NVIDIA GPU Operator, vLLM Production Stack, and Xinference.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.