Overview
Overview
Run local language-model inference on an Ubuntu 24.04.4 LTS environment with Ollama and Qwen2.5 7B Q4_K_M prepared on the AMI. The configuration is intended for teams that want to evaluate or operate a self-managed inference endpoint on their own Amazon EC2 infrastructure. Included configuration
- Ollama 0.34.1 is configured as a systemd service.
- Qwen2.5 7B Q4_K_M (qwen2.5:7b) is documented as preloaded, so the model does not need to be pulled before initial local inference.
- The documented GPU configuration uses Ollama's bundled GPU runner with the NVIDIA driver; a separate system CUDA Toolkit and nvcc are not part of the documented installation.
- The Ollama API uses the local instance endpoint on port 11434; it has no built-in authentication by default. Documented deployment baseline
Product documentation records Ubuntu 24.04.4 LTS on x86_64 and GPU inference checks with an NVIDIA Tesla T4. GPU availability and inference performance depend on the selected EC2 instance, driver, model, and workload. This is not a claim that every instance type or workload has been tested; verify the final AMI and selected instance before production use. Security and customer responsibilities
Keep the unauthenticated Ollama API private. Do not expose port 11434 directly to untrusted networks or the public internet. For remote access, place the API behind customer-managed authentication and TLS, restrict source addresses with network controls, and monitor access. Customers are responsible for instance selection, IAM and network configuration, data governance, backups, and ongoing software/model updates. Pricing
Marketplace software charges, if approved for publication, apply to the seller-provided integration and prepared deployment configuration described above. Amazon EC2 infrastructure and any separately disclosed third-party charges are billed independently.
Highlights
- Ollama is configured as a systemd service on Ubuntu 24.04.4 LTS.
- Qwen2.5 7B Q4_K_M is documented as preloaded for local inference.
- The documented deployment baseline includes NVIDIA GPU support; verify instance compatibility before launch.
Details
Introducing multi-product solutions
You can now purchase comprehensive solutions tailored to use cases and industries.
Features and programs
Financing for AWS Marketplace purchases
Pricing
Dimension | Cost/hour |
|---|---|
g4dn.xlarge Recommended | $350.00 |
g4dn.2xlarge | $350.00 |
g4dn.4xlarge | $350.00 |
g4dn.8xlarge | $350.00 |
g4dn.12xlarge | $350.00 |
g4dn.16xlarge | $350.00 |
g4dn.metal | $350.00 |
Vendor refund policy
No Refund
How can we make this page better?
Legal
Vendor terms and conditions
Content disclaimer
Delivery details
64-bit (x86) Amazon Machine Image (AMI)
Amazon Machine Image (AMI)
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Version 1.0.0 - Initial Release
This release provides an Ubuntu 24.04.4 LTS environment with Ollama and the Qwen2.5 7B Q4_K_M model preloaded. Ollama is configured as a systemd service. The documented baseline includes NVIDIA GPU support and local API access.
Included software:
- Ollama 0.34.1
- Qwen2.5 7B Q4_K_M (qwen2.5:7b)
- Ubuntu 24.04.4 LTS
Customer responsibilities and limitations:
- Verify the final AMI's software versions and model artifact before use.
- The Ollama API does not provide authentication by default. Do not expose it directly to untrusted networks; configure an authenticated proxy and restrict network access where remote access is required.
Additional details
Usage instructions
Launch a compatible GPU-backed EC2 instance from this AMI and select an EC2 key pair. Allow SSH (TCP 22) only from a trusted administrator CIDR. Connect as ubuntu:
ssh -i <key-path> ubuntu@<instance-ip>Verify the GPU, Ollama service, and preloaded model:
nvidia-smi ollama --version sudo systemctl is-active ollama ollama listStart an interactive session with the preloaded model:
ollama run qwen2.5:7bThe local Ollama API listens on port 11434 and does not provide authentication by default. Test it from the instance:
curl <http://localhost:11434/api/generate> -d '{ "model": "qwen2.5:7b", "prompt": "Return the word ready.", "stream": false }'Do not expose port 11434 to the public internet. Keep it local, or configure customer-managed authentication/TLS and restrict network access before enabling remote clients. For troubleshooting, inspect the service with sudo systemctl status ollama and sudo journalctl -u ollama -n 100 --no-pager.
Support
Vendor support
If you encounter problems in the process of using the system, please feel free to contact us by email: support@thinkclouds.ai . Thank you!
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.