Built on openSUSE Linux, this product provides private AI using the Qwen 3 model with 14 billion parameters. MultiCortex HPC (High-Performance Computing) allows you to boost your AI's response quality. This is a plug-and-play, low-cost product with no token fees.
Qwen is the large language model and large multimodal model series of the Qwen Team, Alibaba Group. Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Uniquely support of seamless switching between thinking mode (for complex logical reasoning, math, and coding) and non-thinking mode (for efficient, general-purpose dialogue) within single model, ensuring optimal performance across various scenarios.
Significantly enhancement in its reasoning capabilities, surpassing previous QwQ (in thinking mode) and Qwen2.5 instruct models (in non-thinking mode) on mathematics, code generation, and commonsense logical reasoning.
Superior human preference alignment, excelling in creative writing, role-playing, multi-turn dialogues, and instruction following, to deliver a more natural, engaging, and immersive conversational experience.
Expertise in agent capabilities, enabling precise integration with external tools in both thinking and unthinking modes and achieving leading performance among open-source models in complex agent-based tasks.
Support of 100+ languages and dialects with strong capabilities for multilingual instruction following and translation.
Qwen 3-14b has the following features:
Type: Causal Language Models
Training Stage: Pretraining & Post-training
Number of Parameters: 14.8B
Number of Paramaters (Non-Embedding): 13.2B
Number of Layers: 40
Number of Attention Heads (GQA): 40 for Q and 8 for KV
Context Length: 32,768 natively and 131,072 tokens with YaRN.
Highlights
Boost your private AI system with MultiCortex HPC and leave the thousands of AI updates to us
Enjoy full technical compliance with complete data control in the hands of your company
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 31 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
If you are an AWS Free Tier customer with a free plan, you are eligible to subscribe to this offer. You can use free credits to cover the cost of eligible AWS infrastructure. See AWS Free Tier for more details. If you created an AWS account before July 15th, 2025, and qualify for the Legacy AWS Free Tier, Amazon EC2 charges for Micro instances are free for up to 750 hours per month. See Legacy AWS Free Tier for more details.
You pay by the hour for the EC2 instance type you run. Each dimension maps to a specific AWS instance size, so pricing scales with the compute you choose. Options span general-purpose, compute-optimized, memory-optimized, storage, and GPU or accelerator families, from small shared instances to large bare-metal and high-performance configurations. You pick the instance that fits your workload and pay only for the hours it runs. There are no token fees; the software runs privately in your own environment. Larger or GPU-backed instances carry higher hourly rates than smaller general-purpose ones.
Top-of-mind questions for buyers
What does one hourly unit cover, and what am I actually paying for?
Each unit is one hour of running a specific EC2 instance type with the software installed. You pay for the compute size you pick plus the software. There are no token charges. The model runs privately in your own environment, so usage stays local.
Am I charged when the instance is stopped or paused?
Software charges meter running instance-hours only. A fully stopped instance stops accruing software charges. Underlying AWS storage or reserved resources may still bill separately through AWS, but the hourly software rate applies to running time. Stopping the instance halts the per-hour software cost.
Why do GPU instance types cost more per hour than general-purpose ones?
Each dimension maps to a distinct EC2 instance family and size. GPU and accelerator families carry more compute and hardware, so their hourly rate is higher. General-purpose and smaller instances run at lower hourly rates. You match the instance to your workload and pay the rate tied to that choice.
www.multicortex.ai
Helpful?
Vendor refund policy
If you need to request a refund for software sold by Amazon Web Services, LLC, please contact AWS Customer Service.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Qwen3 model and all packages installed automatically.
Additional details
Usage instructions
Additional details Usage instructions
<br><br>
To use the product, follow these steps: 1. Open a web browser and go to the application at the following address: https://<EC2_Instance_Public_DNS>/index.html. 2. Log in using the credentials below: - Username: ec2-user - Password: the instance_id of your instance. After launching an EC2 instance, wait about 15 minutes for the system to complete the automatic configuration and download of the LLM template. After this time, you can access the chat interface by typing the instance's IP address followed by port 7000 into your browser. For example: http://10.21.103:7000.
Support
Vendor support
No token fees, enhanced performance, and lower consumption of computing resources
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
This product has charges associated with it for support from the seller. Built on openSUSE Linux, this product provides private AI using the Qwen 2.5 model with 0.5 billion parameters. MultiCortex HPC (High-Performance Computing) allows you to boost your AI's response quality. This is a plug-and-play, low-cost product with no token fees.
This deployment package enables seamless hosting of the Qwen/Qwen3-14B language model on Intel® Xeon® processors using the VLLM CPU-optimized Docker image. Designed for efficient inference on CPU-only environments, this solution leverages vLLM lightweight architecture to deliver fast and scalable performance without requiring GPU acceleration. Ideal for enterprise-grade NLP tasks, it offers a cost-effective and accessible way to run large language models on Intel-powered infrastructure.
A self-hosted, production-ready Qwen 3.6 35B model, with 3 billion active parameters deployed into your AWS environment with a single click. Because everything runs entirely within your private cloud, your data stays secure, isolated, and fully under your control. Best of all, unlimited tokens.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.