Amazon Bedrock announces up to 80% lower prices for OpenAI GPT‑5.6 models
Luna drops 80% and Terra drops 20%—your high-volume AI workloads on Bedrock just got dramatically cheaper, automatically.
View original announcement →Visual Summary
What's New
Effective July 30, 2026, Amazon Bedrock has reduced on-demand inference prices for OpenAI GPT‑5.6 Luna by 80% and GPT‑5.6 Terra by 20%, mirroring OpenAI's own first-party pricing changes. These reductions apply automatically with no configuration changes required from customers. GPT‑5.6 Sol pricing remains unchanged, and all three models continue to be accessible via the OpenAI Responses API on the bedrock-mantle endpoint.
How It Works
- GPT‑5.6 Luna and Terra are served through the OpenAI Responses API on the
bedrock-mantleendpoint within Amazon Bedrock, allowing customers to use a familiar API surface without managing model infrastructure. - Price reductions are applied automatically at the billing layer; existing API calls, SDKs, and integrations require no code or configuration changes to benefit from the new rates.
- GPT‑5.6 Luna is architected for fast, high-throughput inference and supports tool use, enabling it to orchestrate multi-step agentic workflows at scale.
- GPT‑5.6 Terra is positioned as a mid-tier model that balances reasoning capability, latency, and cost, targeting production workloads that require more sophisticated outputs than Luna but at lower cost than Sol.
- GPT‑5.6 Sol, the highest-capability model in the family (frontier reasoning, advanced agentic performance), retains its existing pricing and is unaffected by this announcement.
- All three GPT‑5.6 models are listed alongside other OpenAI offerings on Bedrock, including GPT‑5.5, GPT‑5.4, and open-source safeguard/general-purpose models (20B and 120B variants).
Why It's Important
- An 80% price reduction for Luna dramatically lowers the unit economics of high-volume AI workloads such as document classification, content moderation, and customer-service automation, making previously cost-prohibitive scale now financially viable.
- The 20% reduction for Terra improves the cost-performance ratio for everyday production workloads that require reasoning beyond simple classification, broadening its applicability across enterprise use cases.
- Automatic price application means customers immediately realize savings without engineering effort, reducing operational overhead and time-to-value.
- Lower per-token costs incentivize customers to process larger datasets, run longer context windows, or increase inference frequency—unlocking new product capabilities that were previously gated by budget constraints.
- The parity with OpenAI's first-party pricing signals that Amazon Bedrock is committed to competitive, transparent pricing for third-party models, reinforcing it as a credible multi-model platform rather than a premium reseller.
How It's Different
- Unlike direct OpenAI API access, Amazon Bedrock provides these models within AWS's security and compliance boundary, enabling customers to leverage IAM, VPC endpoints, AWS CloudTrail, and AWS PrivateLink without additional integration work.
- Bedrock's unified API and model catalog allow teams to switch between OpenAI, Anthropic, Amazon Nova, and other providers using consistent tooling, reducing vendor lock-in compared to using OpenAI's platform exclusively.
- The
bedrock-mantleendpoint abstracts the underlying OpenAI Responses API, meaning customers can apply Bedrock-native features such as Guardrails, Model Evaluation, and Intelligent Prompt Routing on top of GPT‑5.6 models. - Pricing parity with OpenAI's first-party rates eliminates the traditional cost premium associated with accessing third-party models through a cloud marketplace, making the AWS integration cost-neutral relative to direct API usage.
- The three-tier GPT‑5.6 family (Luna/Terra/Sol) on Bedrock provides a structured cost-performance ladder within a single provider family, giving architects clear upgrade/downgrade paths without switching ecosystems.
When to Prefer It
- Choose GPT‑5.6 Luna for high-volume, latency-sensitive pipelines such as real-time content tagging, bulk email classification, customer-service chatbot responses, or any workload where throughput and cost per call are the primary constraints.
- Choose GPT‑5.6 Luna when building multi-step agentic workflows that invoke tools repeatedly, where the 80% price cut makes iterative tool-calling loops economically feasible at scale.
- Choose GPT‑5.6 Terra for production applications requiring nuanced reasoning—such as summarization of complex documents, code review assistance, or structured data extraction—where Luna's capability ceiling is insufficient but Sol's cost is unjustifiable.
- Choose GPT‑5.6 Sol (unchanged pricing) when the task demands frontier-level reasoning, advanced coding, cybersecurity analysis, or scientific research where output quality is the overriding concern and cost is secondary.
- Prefer this Bedrock-hosted option over direct OpenAI API access when your organization requires AWS-native security controls, consolidated billing, audit logging via CloudTrail, or compliance with data residency requirements within supported AWS regions.
- Consider these models for cost-optimization refactoring of existing workloads currently running on more expensive models (e.g., GPT‑5.5 or Sol), where the new Luna/Terra pricing may deliver acceptable quality at a fraction of the cost.
Availability
- Status: Generally Available (GA) as of July 30, 2026; no preview or waitlist indicated.
- Supported Regions: US East (N. Virginia), US East (Ohio), and US West (Oregon) only; no European or Asia-Pacific regions announced at this time.
- API Endpoint: Accessible via the OpenAI Responses API on the
bedrock-mantleendpoint within Amazon Bedrock. - Pricing Model: On-demand inference; GPT‑5.6 Luna reduced by 80%, GPT‑5.6 Terra reduced by 20% from prior rates; exact per-token prices available on the Amazon Bedrock pricing page.
- Automatic Application: No customer action required—new prices apply automatically to all existing and new API calls from July 30, 2026.
- Limitations: GPT‑5.6 Sol pricing is unchanged; batch inference pricing for GPT‑5.6 models is not mentioned in this announcement; regional availability is currently limited to three US regions.