OpenAI GPT-5.6 Luna is the fastest and most cost-efficient model in the GPT-5.6 family. It is designed for teams building everyday AI workloads where responsiveness, scale, and practical performance matter.
Luna is well suited for workflow automation, content and knowledge tasks, lightweight coding assistance, and other common use cases that benefit from fast, dependable model performance. It gives developers a practical way to bring GPT-5.6 capabilities into production workloads that need to run efficiently.
Through Amazon Bedrock, organizations can use GPT-5.6 Luna in the AWS environment where many enterprise applications and operational workflows already run. Teams can access OpenAI model capabilities alongside their existing AWS security, governance, procurement, billing, and operational workflows.
Highlights
A fast, cost-efficient GPT-5.6 model for everyday AI workflows and applications where responsiveness and efficiency matter
Designed for teams that need strong capability while optimizing for speed, scale, and cost.
Available through Amazon Bedrock so customers can deploy OpenAI model capabilities within their existing AWS environment.
AWS Marketplace now accepts line of credit payments through the PNC Vendor Finance program. This program is available to select AWS customers in the US, excluding NV, NC, ND, TN, & VT.
You pay per token, with no upfront commitment. Charges split into input tokens, output tokens, and cache-related tokens (cached input, cache writes, and cache reads). Each of these applies across processing modes: standard, priority, flex, and batch. You also choose a context scope, either regular or long-context, and a deployment scope, either regional or global. A separate set of cache-write dimensions covers a 30-minute retention window. This structure lets you match each request to a speed, context length, and cost combination, so pricing scales with how you route work.
Top-of-mind questions for buyers
What counts as one billable unit across these token dimensions?
Each unit is one token processed by the model. Input tokens cover text you send. Output tokens cover text the model generates. Cache-related tokens count tokens read from or written to a reuse cache. You are billed per token consumed in each category.
How do the different processing modes change what I pay?
Standard is the default speed. Priority processing runs faster requests. Flex processing favors lower-cost, less time-sensitive work. Batch processing handles grouped requests. Each mode has its own per-token rate across input, output, and cache dimensions, so your route choice sets the price.
Which token type usually drives most of my bill?
Output tokens and input tokens both meter independently and appear together on your invoice. Output tokens often cost more per token than input. Cache reads reduce repeated input costs when prompts reuse content. Cache writes add a charge to store that content, with a 30-minute retention option.
openai.com
Helpful?
Vendor refund policy
All sales are final. Fees are non-refundable.
How can we make this page better?
Tell us how we can improve this page, or report an issue with this product.
Give us feedbackReport a problem with this product or seller
Legal
Vendor terms and conditions
Upon subscribing to this product, you must acknowledge and agree to the terms and conditions outlined in the vendor's End User License Agreement (EULA).
Content disclaimer
Vendors are responsible for their product descriptions and other product content. AWS does not warrant that vendors' product descriptions or other product content are accurate, complete, reliable, current, or error-free.
SaaS delivers cloud-based software applications directly to customers over the internet. You can access these applications through a subscription model. You will pay recurring monthly usage fees through your AWS bill, while AWS handles deployment and infrastructure management, ensuring scalability, reliability, and seamless integration with other AWS services.
AWS Support is a one-on-one, fast-response support channel that is staffed 24x7x365 with experienced and technical support engineers. The service helps customers of all sizes and technical abilities to successfully utilize the products and features provided by Amazon Web Services.
OpenAI GPT-5.5 brings OpenAIs most capable model for complex professional work to coding, analysis, software operation,and long-running agentic tasks through OpenAI APIs and Amazon Bedrock.
OpenAI GPT-5.6 Sol brings OpenAI flagship GPT-5.6 model for advanced reasoning, coding, scientific research, cybersecurity, and agentic workflows to Amazon Bedrock.
OpenAI GPT-5.6 Terra brings a balanced GPT-5.6 model for everyday work, software engineering, knowledge workflows, and scalable AI applications to Amazon Bedrock.
Be the first to review this product. We've partnered with PeerSpot to gather customer feedback. You can share your experience by writing or recording a review, or scheduling a call with a PeerSpot analyst.