Overview
AI Inference Gateway and Model Router by Code Creator
Deploy a private, self-hosted AI gateway on AWS with browser-based AI chat and an OpenAI-compatible API - without building the infrastructure from scratch. AI Inference Gateway and Model Router by Code Creator delivers a fully integrated stack that is ready to serve your team from first boot.
Who This Is For
This product is built for developers, AI teams, small businesses, and internal enterprise users who need a simple, private AI launch point on AWS. Whether you need a browser-based chat interface for internal users or an OpenAI-compatible API endpoint for applications and workflows, this AMI provides a turnkey starting point.
Example scenario: A development team that wants to test prompt chains and AI workflows against a local model before deploying to production can use the OpenAI-compatible API gateway to iterate privately, then route requests to external LLM providers when ready - all through the same unified API interface.
What Is Included
- Open WebUI - Browser-based chat interface supporting Ollama and OpenAI-compatible APIs
- Ollama - Local model serving with a CPU-friendly default model pre-configured and ready to use
- LiteLLM - Unified model routing proxy providing an OpenAI-compatible API layer across many LLM providers
- PostgreSQL - Integrated database for persistent data storage
- Redis - In-memory data store for caching and session management
- Nginx - Reverse proxy handling web traffic to the UI and API endpoints
- First Boot Automation - Automated provisioning that configures the full stack and sets up your instance URL on launch
- Helper Commands - Operational tooling for managing services, checking status, and common administrative tasks
How It Works
Open WebUI provides the browser interface where users interact with AI models through chat. Ollama serves open models locally on the instance. LiteLLM acts as the model routing and API gateway layer, providing a unified OpenAI-compatible interface that can proxy requests to many LLM providers. Nginx sits in front of the stack as a reverse proxy, routing traffic to the appropriate services.
At first boot, the automated provisioning process configures all components, sets up the instance URL, and prepares the CPU-friendly default local model so you can start chatting and making API calls immediately.
Key Capabilities
- Private deployment - Your AI inference traffic stays on your own AWS infrastructure
- Browser-based AI chat - Give your team an intuitive chat interface without external dependencies
- OpenAI-compatible API - Connect downstream applications, scripts, and workflows through a standard API
- Multi-provider model routing - Use LiteLLM to route requests across different LLM providers through a single endpoint
- Local model serving - Run open models directly on the instance using Ollama
- Integrated data layer - PostgreSQL and Redis are pre-configured and running as part of the stack
- Operational tooling - Built-in helper commands simplify day-to-day management
Getting Started
Launch the AMI on your preferred EC2 instance. The first boot automation handles provisioning and configuration of all stack components. Once complete, access Open WebUI through your browser to start chatting with the default local model, or connect to the LiteLLM API endpoint to integrate with your applications. Add external LLM provider API keys through LiteLLM to route requests to additional models as needed.
What You Are Paying For
Code Creator charges apply for the integrated AI inference and model routing stack, automated first boot provisioning, local model serving configuration, OpenAI-compatible API gateway setup, PostgreSQL and Redis integration, operational helper tooling, and AWS Marketplace ready AMI engineering. Standard AWS infrastructure charges for your EC2 instance apply separately.
Note: This product packages and integrates open-source components into a cohesive, pre-configured deployment on AWS. The value is in the integration, automation, and operational readiness - not the individual software components themselves.
Highlights
- Private AI Chat and Local Model Serving Without Building the Stack Yourself Launch a self-hosted AI workspace with Open WebUI for browser-based chat, Ollama for local model serving, and a CPU-friendly default model - all pre-configured and ready to use after first boot. Keep your AI conversations and data on your own AWS infrastructure instead of routing them through third-party services.
- OpenAI-Compatible API Gateway for Developers and Internal Workflows LiteLLM provides a unified OpenAI-compatible API endpoint that your developers, applications, and internal AI workflows can call directly. Route requests to your local model or connect external LLM providers through a single gateway backed by PostgreSQL and Redis for reliable operation.
- Automated AWS Deployment with First-Boot Configuration The AMI includes automated first-boot public IP detection, Nginx reverse proxy setup, and helper commands so you spend less time on infrastructure plumbing. Six open-source components - Open WebUI, Ollama, LiteLLM, PostgreSQL, Redis, and Nginx - are pre-integrated and configured to work together on Ubuntu 24.04.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Cost/hour |
|---|---|
t3.xlarge Recommended | $0.035 |
m6i.32xlarge | $0.035 |
m6i.24xlarge | $0.035 |
m7i.12xlarge | $0.035 |
m6i.12xlarge | $0.035 |
m7i.24xlarge | $0.035 |
m7i.16xlarge | $0.035 |
m6i.16xlarge | $0.035 |
m7i.48xlarge | $0.035 |
m6i.2xlarge | $0.035 |
Vendor refund policy
No contracts. We do not currently support refunds, but you can cancel at any time.
Custom pricing options
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
64-bit (x86) Amazon Machine Image (AMI)
Amazon Machine Image (AMI)
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Version 2.6 includes Open WebUI, Ollama, LiteLLM, PostgreSQL, Redis, Nginx, first boot public IP refresh, automatic Docker stack startup, helper commands, a Code Creator help page, API information page, and the default llama3.2 1b CPU friendly model.
Additional details
Resources
Vendor resources
Support
Vendor support
For support inquiries related to the AI Inference Gateway and Model Router AMI, contact Code Creator by email at info@codecreator.com .
The following open-source documentation resources are available for the individual components included in this AMI:
- Open WebUI Documentation: https://docs.openwebui.com/
- LiteLLM Documentation: https://docs.litellm.ai/docs/
- Ollama Quickstart Guide:
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.