Listing Thumbnail

    Hardened vLLM on Ubuntu 24.04 LTS

     Info
    Sold by: EASYCLOUD 
    Deployed on AWS
    AWS Free Tier
    This is a repackaged open source software product wherein additional charges apply for a hardened, AWS-optimized image with verification reports and maintenance. This AMI delivers vLLM 0.24 as a bring-your-own-model OpenAI-compatible inference server, on an Ubuntu 24.04 LTS base tuned for EC2 by Easycloud-with a provenance evidence pack and SBOM.

    Overview

    This is a repackaged open source software product wherein additional charges apply for a hardened, AWS-optimized image with verifiable provenance evidence and maintenance. vLLM 0.24 runs on an Ubuntu 24.04 LTS base tuned for Amazon EC2, with a documented set of security settings applied on top and a complete provenance evidence pack including a Software Bill of Materials (SBOM).

    Value-Added Features

    vLLM 0.24 Deployment:

    1. Bring Your Own Model: No model is pre-installed. One built-in command downloads and activates any Hugging Face model you choose, so nothing has to be deleted and no disk is wasted on weights you will not serve.

    2. High-Throughput OpenAI-Compatible Serving: PagedAttention memory management and continuous batching maximise tokens per second, behind a drop-in OpenAI-compatible REST API that existing SDKs and clients use unchanged.

    3. Cold-Disk Prewarm: A boot-time step reads the engine's startup file set into page cache, so service starts stay fast on a freshly launched instance instead of crawling through lazily loaded EBS blocks.

    Comprehensive Evidence Pack:

    Every AMI carries a provenance evidence pack under /opt/optimization-report/:

    • Tuning_Parameters.txt: Every tuning change, shown against its base-image value, with the reasoning and how to reverse it.
    • Security_Parameters.txt: All 31 security settings in the same format; the 10 already at the wanted value are marked unchanged rather than presented as improvements.
    • package_changes.txt: Packages upgraded, installed and removed versus the base image.
    • sbom.spdx.json: Software Bill of Materials in SPDX format, generated from the dpkg database.
    • cve_scan.txt / cve_scan_full.txt.gz: Vulnerability scan grouped by whether an upstream fix exists, with the GPU stack pinning explained.
    • key_files.sha256: SHA-256 checksums of all 7 files this image changed, so the delivered state can be verified in one command.
    • README.txt: Index of artifact contents and operational guidance.

    Security & Attack Surface:

    1. Kernel Network Hardening: 19 network-related kernel parameters are fixed in /etc/sysctl.d/99-security.conf: ICMP redirects are neither accepted nor sent, source-routed packets are rejected, and the rest are pinned at safe values so they cannot drift.

    2. Unused Kernel Modules Disabled: 12 rarely used filesystem and network protocol modules are blocked from loading. No NVIDIA, EFA or container module is affected.

    3. SSH Login Window Tightened: LoginGraceTime is reduced from 120 to 60 seconds. Root login policy, authentication retries and forwarding are left as the base image ships them.

    4. No Outstanding Upstream Fixes: Every package outside the GPU pins was upgraded to the newest version its distribution offers, including updates Ubuntu was still phasing in. The scan reports zero findings with a fix available and not installed.

    GPU Stack and Engine Left Intact:

    1. Nothing Touches the Stack: The NVIDIA driver, CUDA runtime, vLLM virtual environment and its three systemd units are not modified. After a reboot nvidia-smi reports the same driver, the engine imports and runs a GPU operation, and the kernel command line is byte for byte identical to the base image - verified on a GPU instance.

    2. Version Pinning Respected: The kernel and GPU stack are held at fixed versions because the driver is built against one specific kernel. That pinning predates this release and is documented in the evidence pack rather than silently overridden.

    AWS Network & Kernel Tuning:

    1. Network Stack: TCP BBR congestion control with fair queueing, increased connection backlogs, raised socket buffer ceilings, and MTU probing enabled. Ceilings are only ever raised, never lowered.

    2. System Tuning: tuned daemon active with an aws-optimized profile whose bootloader section is deliberately empty, so the network interface keeps its original name after a reboot.

    3. Reliability: The systemd journal is capped at 200 MB so logs cannot fill the root filesystem.

    Operational Tools:

    1. Preinstalled Tools: AWS SSM Agent active, CloudWatch Agent installed (disabled by default).

    Maintenance:

    1. Maintenance Specifications: Rebuilt and updated bi-weekly to monthly incorporating upstream Ubuntu 24.04 security updates, with artifact documentation refreshed per release.

    About vLLM

    A leading open-source LLM serving engine built around PagedAttention and continuous batching for high-throughput GPU inference. Running it on your own instance keeps prompts and completions inside your AWS account, and it pairs with front-ends such as Open WebUI, LibreChat, Dify and Langflow.

    About Ubuntu 24.04 LTS

    The most widely deployed Linux distribution in the cloud, built on the Debian foundation, with a five-year maintenance window for LTS releases.

    Highlights

    • vLLM 0.24 on an AWS-Optimized Ubuntu 24.04 Base: 19 security-related kernel parameters, 12 blocked modules and a tightened SSH login window applied on top, with the driver, CUDA runtime, engine environment and kernel command line left untouched.
    • Verifiable Evidence Pack and SBOM: An 8-file evidence pack ships inside the image, with an SPDX Software Bill of Materials, a CVE scan grouped by upstream fix availability, a package delta report, and SHA-256 checksums of every file the image changed.
    • vLLM 0.24: Best-practice deployment. Bring your own model with one built-in command, served through a drop-in OpenAI-compatible API with PagedAttention and continuous batching, and a boot-time prewarm for fast starts on fresh instances.

    Details

    Delivery method

    Delivery option
    64-bit (x86) Amazon Machine Image (AMI)

    Latest version

    Operating system
    Ubuntu 24.04

    Deployed on AWS
    New

    Introducing multi-product solutions

    You can now purchase comprehensive solutions tailored to use cases and industries.

    Multi-product solutions

    Features and programs

    Financing for AWS Marketplace purchases

    AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
    Financing for AWS Marketplace purchases

    Pricing

    Hardened vLLM on Ubuntu 24.04 LTS

     Info
    Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time. Alternatively, you can pay upfront for a contract, which typically covers your anticipated usage for the contract duration. Any usage beyond contract will incur additional usage-based costs.
    Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator  to estimate your infrastructure costs.
    If you are an AWS Free Tier customer with a free plan, you are eligible to subscribe to this offer. You can use free credits to cover the cost of eligible AWS infrastructure. See AWS Free Tier  for more details. If you created an AWS account before July 15th, 2025, and qualify for the Legacy AWS Free Tier, Amazon EC2 charges for Micro instances are free for up to 750 hours per month. See Legacy AWS Free Tier  for more details.

    Usage costs (791)

     Info
    • ...
    Dimension
    Cost/hour
    g4dn.xlarge
    Recommended
    $0.19
    t2.micro
    $0.03
    t3.micro
    $0.03
    t3.small
    $0.04
    m5ad.8xlarge
    $0.39
    g6.4xlarge
    $0.29
    z1d.6xlarge
    $0.39
    x8aedz.12xlarge
    $0.68
    d3.4xlarge
    $0.29
    g2.8xlarge
    $0.39

    AI Insights

     Info

    Dimensions summary

    You pay by the hour for the EC2 instance type you run this hardened vLLM software on. Each dimension maps to a specific AWS instance size, so pricing scales with the compute you choose. Smaller general-purpose instances cost less per hour, while larger memory-, compute-, storage-, and GPU-optimized instances cost more. You are billed only for the hours each instance runs. Pick the instance that fits your model-serving workload and budget. There are no upfront commitments or fixed terms; billing follows your actual usage.

    Top-of-mind questions for buyers

    The rate covers the hardened vLLM software running on that specific EC2 instance size for each hour it runs. You pick one instance type per deployment, and its per-hour price reflects the compute, memory, storage, or GPU resources that size provides. AWS infrastructure charges are separate.
    Software charges accrue only while the instance runs. A stopped or paused instance stops the hourly software meter. You may still owe underlying AWS storage fees for attached volumes, but the vLLM software bills on running hours only.
    Yes. The software follows security hardening practices such as CIS Benchmark alignment, firewall configuration, and access controls regardless of which instance size you run. The instance type you pick changes only the compute resources and the hourly price, not the hardening applied.
    www.easyclouds.io
    Helpful?

    Vendor refund policy

    No refunds. Cancel anytime.

    How can we make this page better?

    Tell us how we can improve this page, or report an issue with this product.
    Tell us how we can improve this page, or report an issue with this product.

    Legal

    Vendor terms and conditions

    Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA) .

    Content disclaimer

    Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.

    Usage information

     Info

    Delivery details

    64-bit (x86) Amazon Machine Image (AMI)

    Amazon Machine Image (AMI)

    An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.

    Version release notes

    Latest Updates

    Additional details

    Usage instructions

    Connection Methods

    Once launched, SSH into the instance. The default username is 'ubuntu'. You can switch to the root user environment by running: sudo su -

    This product requires SSH for initial setup: you must log in once to choose and download a model before the API port starts serving.

    Install Information

    • OS: Ubuntu 24.04 LTS (x86_64, Minimal Installation)

    • vLLM: 0.24 (OpenAI-compatible inference server, installed at /opt/vllm)

    • Runtime: Python virtual environment with CUDA GPU support

    • Config File: '/etc/vllm/vllm.env' (written by the vllm-model helper; do not edit by hand)

    • No model is pre-installed (bring your own model): models are downloaded to /opt/models/<model-name> after you choose one

    Usage Instructions

    1. Launch on a GPU instance. Recommended: g4dn.xlarge (16 GB GPU memory, runs models up to about 4B parameters) or larger g5/g6 instances for bigger models. Match your model choice to the GPU memory of your instance - a model that is too large will fail to start.

    2. The instance boots in under 3 minutes. At this point the vllm service is intentionally inactive and port 8000 is not yet listening - the login banner shows NO MODEL CONFIGURED YET. This is the normal factory state, not a fault.

    3. SSH in and run one command to download and activate your chosen model - time varies by model size: sudo vllm-model Qwen/Qwen3-4B-Instruct-2507 This example fits 16 GB GPUs; any Hugging Face model ID is supported. Run vllm-model with no arguments to see more examples per GPU memory tier.

    4. Note on model downloads: the instance downloads model files from the Hugging Face Hub, which requires outbound internet access. Model sizes range from hundreds of MB to tens of GB depending on your choice. Without an HF_TOKEN the download is anonymous and subject to Hugging Face rate limits; the vllm-model help output explains how to configure HF_TOKEN.

    5. Retrieve your API key: the key equals this instance's EC2 Instance ID (visible in the AWS console). You can also confirm it on the instance by running: cat /etc/vllm/api-key

    6. Call the OpenAI-compatible API from your client with the header Authorization: Bearer <YOUR_INSTANCE_ID>, for example: curl http://<YOUR_IP>:8000/v1/models -H 'Authorization: Bearer <YOUR_INSTANCE_ID>' Ready-to-copy curl examples are shown in the login banner and in /home/ubuntu/readme-vllm.

    7. Manage the service: sudo systemctl start|stop|restart|status vllm. View logs with: sudo journalctl -u vllm

    8. After an instance stop/start, your downloaded model is preserved and the service resumes automatically, but the inference engine needs about 2-3 minutes to re-initialize before it serves requests again.

    Firewall Configuration

    • SSH (Port 22): Required - this product needs SSH for the one-time initial model setup and ongoing administration.

    • vLLM OpenAI-compatible API (Port 8000): Serves the inference API once a model is configured; restrict to trusted IPs.

    • Security Recommendation: For production environments, strictly limit access to these ports to trusted IP addresses only via cloud Security Groups or the local firewall.

    Support

    Vendor support

    Should you encounter any issues while using the system, please do not hesitate to contact us via email at: support@easyclouds.io ,Thank you!

    AWS infrastructure support

    AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.

    Similar products

    Customer reviews

    Ratings and reviews

     Info
    0 ratings
    5 star
    4 star
    3 star
    2 star
    1 star
    0%
    0%
    0%
    0%
    0%
    0 reviews
    No customer reviews yet
    Be the first to review this product . We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.