This is a repackaged open source software product wherein additional charges apply for a hardened, AWS-optimized image with verification reports and maintenance. This AMI delivers vLLM 0.24 as a bring-your-own-model OpenAI-compatible inference server, on an Ubuntu 24.04 LTS base tuned for EC2 by Easycloud-with a provenance evidence pack and SBOM.
This is a repackaged open source software product wherein additional charges apply for a hardened, AWS-optimized image with verifiable provenance evidence and maintenance. vLLM 0.24 runs on an Ubuntu 24.04 LTS base tuned for Amazon EC2, with a documented set of security settings applied on top and a complete provenance evidence pack including a Software Bill of Materials (SBOM).
Value-Added Features
vLLM 0.24 Deployment:
Bring Your Own Model: No model is pre-installed. One built-in command downloads and activates any Hugging Face model you choose, so nothing has to be deleted and no disk is wasted on weights you will not serve.
High-Throughput OpenAI-Compatible Serving: PagedAttention memory management and continuous batching maximise tokens per second, behind a drop-in OpenAI-compatible REST API that existing SDKs and clients use unchanged.
Cold-Disk Prewarm: A boot-time step reads the engine's startup file set into page cache, so service starts stay fast on a freshly launched instance instead of crawling through lazily loaded EBS blocks.
Comprehensive Evidence Pack:
Every AMI carries a provenance evidence pack under /opt/optimization-report/:
Tuning_Parameters.txt: Every tuning change, shown against its base-image value, with the reasoning and how to reverse it.
Security_Parameters.txt: All 31 security settings in the same format; the 10 already at the wanted value are marked unchanged rather than presented as improvements.
package_changes.txt: Packages upgraded, installed and removed versus the base image.
sbom.spdx.json: Software Bill of Materials in SPDX format, generated from the dpkg database.
cve_scan.txt / cve_scan_full.txt.gz: Vulnerability scan grouped by whether an upstream fix exists, with the GPU stack pinning explained.
key_files.sha256: SHA-256 checksums of all 7 files this image changed, so the delivered state can be verified in one command.
README.txt: Index of artifact contents and operational guidance.
Security & Attack Surface:
Kernel Network Hardening: 19 network-related kernel parameters are fixed in /etc/sysctl.d/99-security.conf: ICMP redirects are neither accepted nor sent, source-routed packets are rejected, and the rest are pinned at safe values so they cannot drift.
Unused Kernel Modules Disabled: 12 rarely used filesystem and network protocol modules are blocked from loading. No NVIDIA, EFA or container module is affected.
SSH Login Window Tightened: LoginGraceTime is reduced from 120 to 60 seconds. Root login policy, authentication retries and forwarding are left as the base image ships them.
No Outstanding Upstream Fixes: Every package outside the GPU pins was upgraded to the newest version its distribution offers, including updates Ubuntu was still phasing in. The scan reports zero findings with a fix available and not installed.
GPU Stack and Engine Left Intact:
Nothing Touches the Stack: The NVIDIA driver, CUDA runtime, vLLM virtual environment and its three systemd units are not modified. After a reboot nvidia-smi reports the same driver, the engine imports and runs a GPU operation, and the kernel command line is byte for byte identical to the base image - verified on a GPU instance.
Version Pinning Respected: The kernel and GPU stack are held at fixed versions because the driver is built against one specific kernel. That pinning predates this release and is documented in the evidence pack rather than silently overridden.
AWS Network & Kernel Tuning:
Network Stack: TCP BBR congestion control with fair queueing, increased connection backlogs, raised socket buffer ceilings, and MTU probing enabled. Ceilings are only ever raised, never lowered.
System Tuning: tuned daemon active with an aws-optimized profile whose bootloader section is deliberately empty, so the network interface keeps its original name after a reboot.
Reliability: The systemd journal is capped at 200 MB so logs cannot fill the root filesystem.
Maintenance Specifications: Rebuilt and updated bi-weekly to monthly incorporating upstream Ubuntu 24.04 security updates, with artifact documentation refreshed per release.
About vLLM
A leading open-source LLM serving engine built around PagedAttention and continuous batching for high-throughput GPU inference. Running it on your own instance keeps prompts and completions inside your AWS account, and it pairs with front-ends such as Open WebUI, LibreChat, Dify and Langflow.
About Ubuntu 24.04 LTS
The most widely deployed Linux distribution in the cloud, built on the Debian foundation, with a five-year maintenance window for LTS releases.
Highlights
vLLM 0.24 on an AWS-Optimized Ubuntu 24.04 Base: 19 security-related kernel parameters, 12 blocked modules and a tightened SSH login window applied on top, with the driver, CUDA runtime, engine environment and kernel command line left untouched.
Verifiable Evidence Pack and SBOM: An 8-file evidence pack ships inside the image, with an SPDX Software Bill of Materials, a CVE scan grouped by upstream fix availability, a package delta report, and SHA-256 checksums of every file the image changed.
vLLM 0.24: Best-practice deployment. Bring your own model with one built-in command, served through a drop-in OpenAI-compatible API with PagedAttention and continuous batching, and a boot-time prewarm for fast starts on fresh instances.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time. Alternatively, you can pay upfront for a contract, which typically covers your anticipated usage for the contract duration. Any usage beyond contract will incur additional usage-based costs.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
If you are an AWS Free Tier customer with a free plan, you are eligible to subscribe to this offer. You can use free credits to cover the cost of eligible AWS infrastructure. See AWS Free Tier for more details. If you created an AWS account before July 15th, 2025, and qualify for the Legacy AWS Free Tier, Amazon EC2 charges for Micro instances are free for up to 750 hours per month. See Legacy AWS Free Tier for more details.
You pay by the hour for the EC2 instance type you run this hardened vLLM software on. Each dimension maps to a specific AWS instance size, so pricing scales with the compute you choose. Smaller general-purpose instances cost less per hour, while larger memory-, compute-, storage-, and GPU-optimized instances cost more. You are billed only for the hours each instance runs. Pick the instance that fits your model-serving workload and budget. There are no upfront commitments or fixed terms; billing follows your actual usage.
Top-of-mind questions for buyers
What does the hourly rate cover for each instance type I choose?
The rate covers the hardened vLLM software running on that specific EC2 instance size for each hour it runs. You pick one instance type per deployment, and its per-hour price reflects the compute, memory, storage, or GPU resources that size provides. AWS infrastructure charges are separate.
Am I charged when an instance is stopped or paused?
Software charges accrue only while the instance runs. A stopped or paused instance stops the hourly software meter. You may still owe underlying AWS storage fees for attached volumes, but the vLLM software bills on running hours only.
Does the hardened image apply the same security baseline across every instance type?
Yes. The software follows security hardening practices such as CIS Benchmark alignment, firewall configuration, and access controls regardless of which instance size you run. The instance type you pick changes only the compute resources and the hourly price, not the hardening applied.
www.easyclouds.io
Helpful?
Vendor refund policy
No refunds. Cancel anytime.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Latest Updates
Additional details
Usage instructions
Connection Methods
Once launched, SSH into the instance. The default username is 'ubuntu'. You can switch to the root user environment by running: sudo su -
This product requires SSH for initial setup: you must log in once to choose and download a model before the API port starts serving.
vLLM: 0.24 (OpenAI-compatible inference server, installed at /opt/vllm)
Runtime: Python virtual environment with CUDA GPU support
Config File: '/etc/vllm/vllm.env' (written by the vllm-model helper; do not edit by hand)
No model is pre-installed (bring your own model): models are downloaded to /opt/models/<model-name> after you choose one
Usage Instructions
Launch on a GPU instance. Recommended: g4dn.xlarge (16 GB GPU memory, runs models up to about 4B parameters) or larger g5/g6 instances for bigger models. Match your model choice to the GPU memory of your instance - a model that is too large will fail to start.
The instance boots in under 3 minutes. At this point the vllm service is intentionally inactive and port 8000 is not yet listening - the login banner shows NO MODEL CONFIGURED YET. This is the normal factory state, not a fault.
SSH in and run one command to download and activate your chosen model - time varies by model size: sudo vllm-model Qwen/Qwen3-4B-Instruct-2507
This example fits 16 GB GPUs; any Hugging Face model ID is supported. Run vllm-model with no arguments to see more examples per GPU memory tier.
Note on model downloads: the instance downloads model files from the Hugging Face Hub, which requires outbound internet access. Model sizes range from hundreds of MB to tens of GB depending on your choice. Without an HF_TOKEN the download is anonymous and subject to Hugging Face rate limits; the vllm-model help output explains how to configure HF_TOKEN.
Retrieve your API key: the key equals this instance's EC2 Instance ID (visible in the AWS console). You can also confirm it on the instance by running: cat /etc/vllm/api-key
Call the OpenAI-compatible API from your client with the header Authorization: Bearer <YOUR_INSTANCE_ID>, for example: curl http://<YOUR_IP>:8000/v1/models -H 'Authorization: Bearer <YOUR_INSTANCE_ID>'
Ready-to-copy curl examples are shown in the login banner and in /home/ubuntu/readme-vllm.
After an instance stop/start, your downloaded model is preserved and the service resumes automatically, but the inference engine needs about 2-3 minutes to re-initialize before it serves requests again.
Firewall Configuration
SSH (Port 22): Required - this product needs SSH for the one-time initial model setup and ongoing administration.
vLLM OpenAI-compatible API (Port 8000): Serves the inference API once a model is configured; restrict to trusted IPs.
Security Recommendation: For production environments, strictly limit access to these ports to trusted IP addresses only via cloud Security Groups or the local firewall.
Support
Vendor support
Should you encounter any issues while using the system, please do not hesitate to contact us via email at: support@easyclouds.io,Thank you!
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
This is a repackaged software product wherein additional charges apply for a pre-hardened, SI Core STIG Hardened image and seller support. Ubuntu 24.04 LTS is designed for enterprise environments, providing a robust and secure platform for the deployment of cloud applications. With its long-term support and regular updates, Ubuntu 24.04 LTS ensures stability and reliability for critical workloads. Seamlessly integrate Ubuntu 24.04 LTS into your existing infrastructure, benefiting from its extensive package repository and developer-friendly tools. Ideal for container orchestration, web services, and high-performance computing, Ubuntu 24.04 LTS equips organizations with the flexibility to innovate while maintaining a strong security posture. Choose Ubuntu 24.04 LTS for modern cloud solutions.
The solution goes beyond compliance by offering a Ubuntu server with comprehensive security hardening by default, covering everything from applications to the Linux kernel. With VED threat mitigation, you can rest assured that your digital assets are protected against advanced threats.
This product has charges associated with it for CoreNova hardening, maintenance, validation notes, and seller support. Ubuntu 24.04 LTS Hardened for AWS Graviton (ARM64, LVM, XFS) provides a hardened Ubuntu 24.04 LTS EC2 baseline with SSH lockdown, audit logging, AIDE, firewall controls, and buyer-side OpenSCAP notes.
This is a repackaged software product wherein additional charges apply for a pre-hardened, SI Core STIG Hardened image and seller support. PostgreSQL on Ubuntu 24.04 LTS offers an optimized environment for your relational database needs, ensuring high performance and robust security. With its pre-hardened configuration, PostgreSQL on Ubuntu 24.04 LTS is ideal for enterprises seeking compliance and reliability in sensitive data environments. This AMI simplifies deployment in the AWS EC2 cloud, allowing you to quickly scale your applications. Experience seamless integration, advanced analytics capabilities, and unmatched flexibility with PostgreSQL on Ubuntu 24.04 LTS, designed to support modern workloads and development practices efficiently.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.