Built on openSUSE Linux, this product provides private AI using the ChatGPT OSS 20b model with 20 billion parameters. MultiCortex HPC (High-Performance Computing) allows you to boost your AI's response quality. This is a plug-and-play, low-cost product with no token fees.
gpt-oss-20b is our medium-sized open-weight model for low latency, local, or specialized use-cases (21B parameters with 3.6B active parameters).
Key features
Permissive Apache 2.0 license: Build freely without copyleft restrictions or patent risk ideal for experimentation, customization, and commercial deployment. Configurable reasoning effort: Easily adjust the reasoning effort (low, medium, high) based on your specific use case and latency needs. Full chain-of-thought: Gain complete access to the model's reasoning process, facilitating easier debugging and increased trust in outputs. Fine-tunable: Fully customize models to your specific use case through parameter finetuning. Agentic capabilities: Use the models' native capabilities for function calling, web browsing, Python code execution, and structured outputs.
Highlights
Boost your private AI system with MultiCortex HPC and leave the thousands of AI updates to us
Enjoy full technical compliance with complete data control in the hands of your company
No token fees, enhanced performance, and lower consumption of computing resources
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
Try this product free for 31 days according to the free trial terms set by the vendor. Usage-based pricing is in effect for usage beyond the free trial terms. Your free trial gets automatically converted to a paid subscription when the trial ends, but may be canceled any time before that.
ChatGPT OSS 20b Boosted by openSUSE MultiCortex HPC
Pricing is based on actual usage, with charges varying according to how much you consume. Subscriptions have no end date and may be canceled any time.
Additional AWS infrastructure costs may apply. Use the AWS Pricing Calculator to estimate your infrastructure costs.
If you are an AWS Free Tier customer with a free plan, you are eligible to subscribe to this offer. You can use free credits to cover the cost of eligible AWS infrastructure. See AWS Free Tier for more details. If you created an AWS account before July 15th, 2025, and qualify for the Legacy AWS Free Tier, Amazon EC2 charges for Micro instances are free for up to 750 hours per month. See Legacy AWS Free Tier for more details.
You pay by the hour for the EC2 instance type you run this software on. Each dimension maps to one specific instance size, so pricing scales with the compute you choose. Options span many families: general-purpose, memory-optimized, compute-optimized, storage-optimized, and GPU or accelerator instances. Smaller instances like t3.nano cost less per hour, while large bare-metal and multi-GPU instances like p5.48xlarge cost more. You are not charged per token. Choose the instance that fits your workload size, then run it as long as needed and pay only for the hours used.
Top-of-mind questions for buyers
What does one hourly unit cover, and am I charged per token?
Each hourly unit is one running EC2 instance of the type you pick, billed per hour. You run the ChatGPT OSS 20b software on that instance. There are no per-token fees. You pay only for the instance-hours used while the instance runs.
Am I charged when the instance is stopped or paused?
The software charge meters running time only. A fully stopped instance stops accruing the hourly software fee. Stopped instances may still incur underlying AWS storage costs for attached volumes, but those are separate from this software charge. Start it again and hourly billing resumes.
How do I decide between a GPU instance and a CPU-only instance for this model?
Cost tracks the instance you choose. GPU or accelerator families like p5, g6, and inf2 suit heavier processing. CPU-focused families like m7i, c7i, and r7i suit lighter loads. The heterogeneous design uses available accelerators. Match the instance size to your workload, then pay for the hours you run.
www.multicortex.ai+1
Helpful?
Vendor refund policy
If you need to request a refund for software sold by Amazon Web Services, LLC, please contact AWS Customer Service.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
An AMI is a virtual image that provides the information required to launch an instance. Amazon EC2 (Elastic Compute Cloud) instances are virtual servers on which you can run your applications and workloads, offering varying combinations of CPU, memory, storage, and networking resources. You can launch as many instances from as many different AMIs as you need.
Version release notes
Added gpt-oss-20b for lower latency, and local or specialized use cases (21B parameters with 3.6B active parameters)
Additional details
Usage instructions
To use the product, follow these steps: 1. Open a web browser and go to the application at the following address: https://<EC2_Instance_Public_DNS>/index.html. 2. Log in using the credentials below: - Username: ec2-user - Password: the instance_id of your instance. After launching an EC2 instance, wait about 15 minutes for the system to complete the automatic configuration and download of the LLM template. After this time, you can access the chat interface by typing the instance's IP address followed by port 7000 into your browser. For example: http://10.21.103:7000.
Support
Vendor support
The support service called MultiCortex, developed based on the openSUSE operating system, aims to provide the necessary resources for the efficient execution of AI models that are always up to date and aligned with the latest technological innovations in the area, offering a robust and secure environment for the development and application of AI-based solutions.
AWS infrastructure support
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
LibreChat is an enhanced ChatGPT clone that brings together the latest advancements in AI technology. It serves as a centralized hub for all your AI conversations, providing a familiar, user-friendly interface enriched with advanced features and customization capabilities.
This product has charges associated with it for seller support. By now, everyone agrees that GPT & LLMs are the paradigm shift on how we consume & generate information. At the same time, enterprises & consumers are concerned about the privacy & security of their datasources when interacting with GPTs.
Onyx (formerly Danswer) is the answer to these concerns. It allows you to fully exploit the benefits of GPT to query, chat and analyze your own data without sharing or hosting it with 3rd parties.
This product has charges associated with it for the IronCloud LLM platform and support. Deploy your own secure, ChatGPT-style AI in minutes - fully private, self-hosted, and compliant with NIST 800-171, CMMC, and ITAR. IronCloud LLM runs entirely inside your AWS environment, keeping all data, prompts, and files under your control. No SaaS, no data leaks - just powerful AI you own. Built for regulated industries such as government, defense, finance, and healthcare, IronCloud LLM supports GPT-4, Claude, Gemini, AWS Bedrock, and more.
This product has charges associated with it for hardening, security configuration, and support.
AnythingLLM is a self-hosted ChatGPT-style workspace for chatting with your documents using any LLM provider, with built-in RAG, AI agents and vector search in a single hardened container. Unlike bare AnythingLLM AMIs that ship without TLS, no admin auth, and the server on 0.0.0.0:3001, this Lynxroute build is ready out of the box: admin user with unique password at first boot, server bound to loopback behind Nginx TLS, embedded LanceDB and native embedder pre-configured, on a CIS Level 1 hardened Ubuntu 24.04 LTS base.
MIT license - fully auditable, no vendor lock-in.
This product has a fee associated with the provision and deployment of the application and AMI support. LibreChat gives you the ability to integrate multiple AI models all in one seamless interface platform. It also integrates and enhances original client features such as conversation and message search, prompt templates and plugins.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.