← Back to all announcements
★★★★★ 30/07/2026

Amazon Bedrock announces up to 80% lower prices for OpenAI GPT‑5.6 models

Luna drops 80% and Terra drops 20%—your high-volume AI workloads on Bedrock just got dramatically cheaper, automatically.

View original announcement →

Visual Summary

graph TD A{{GPT-5.6 Price Reduction}}:::announced B(Amazon Bedrock):::compute C([GPT-5.6 Luna -80%]):::feature D([GPT-5.6 Terra -20%]):::feature E([GPT-5.6 Sol Unchanged]):::feature F((Customers)):::external G([OpenAI Responses API]):::feature H([Bedrock-Mantle Endpoint]):::feature F ==>|"invokes"| B B ==>|"serves models"| A A -->|"high-volume tasks"| C A -->|"balanced workloads"| D A -.->|"frontier reasoning"| E F -->|"API calls"| G G -->|"routes via"| H H -->|"connects to"| B classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Effective July 30, 2026, Amazon Bedrock has reduced on-demand inference prices for OpenAI GPT‑5.6 Luna by 80% and GPT‑5.6 Terra by 20%, mirroring OpenAI's own first-party pricing changes. These reductions apply automatically with no configuration changes required from customers. GPT‑5.6 Sol pricing remains unchanged, and all three models continue to be accessible via the OpenAI Responses API on the bedrock-mantle endpoint.

How It Works

  • GPT‑5.6 Luna and Terra are served through the OpenAI Responses API on the bedrock-mantle endpoint within Amazon Bedrock, allowing customers to use a familiar API surface without managing model infrastructure.
  • Price reductions are applied automatically at the billing layer; existing API calls, SDKs, and integrations require no code or configuration changes to benefit from the new rates.
  • GPT‑5.6 Luna is architected for fast, high-throughput inference and supports tool use, enabling it to orchestrate multi-step agentic workflows at scale.
  • GPT‑5.6 Terra is positioned as a mid-tier model that balances reasoning capability, latency, and cost, targeting production workloads that require more sophisticated outputs than Luna but at lower cost than Sol.
  • GPT‑5.6 Sol, the highest-capability model in the family (frontier reasoning, advanced agentic performance), retains its existing pricing and is unaffected by this announcement.
  • All three GPT‑5.6 models are listed alongside other OpenAI offerings on Bedrock, including GPT‑5.5, GPT‑5.4, and open-source safeguard/general-purpose models (20B and 120B variants).

Why It's Important

  • An 80% price reduction for Luna dramatically lowers the unit economics of high-volume AI workloads such as document classification, content moderation, and customer-service automation, making previously cost-prohibitive scale now financially viable.
  • The 20% reduction for Terra improves the cost-performance ratio for everyday production workloads that require reasoning beyond simple classification, broadening its applicability across enterprise use cases.
  • Automatic price application means customers immediately realize savings without engineering effort, reducing operational overhead and time-to-value.
  • Lower per-token costs incentivize customers to process larger datasets, run longer context windows, or increase inference frequency—unlocking new product capabilities that were previously gated by budget constraints.
  • The parity with OpenAI's first-party pricing signals that Amazon Bedrock is committed to competitive, transparent pricing for third-party models, reinforcing it as a credible multi-model platform rather than a premium reseller.

How It's Different

  • Unlike direct OpenAI API access, Amazon Bedrock provides these models within AWS's security and compliance boundary, enabling customers to leverage IAM, VPC endpoints, AWS CloudTrail, and AWS PrivateLink without additional integration work.
  • Bedrock's unified API and model catalog allow teams to switch between OpenAI, Anthropic, Amazon Nova, and other providers using consistent tooling, reducing vendor lock-in compared to using OpenAI's platform exclusively.
  • The bedrock-mantle endpoint abstracts the underlying OpenAI Responses API, meaning customers can apply Bedrock-native features such as Guardrails, Model Evaluation, and Intelligent Prompt Routing on top of GPT‑5.6 models.
  • Pricing parity with OpenAI's first-party rates eliminates the traditional cost premium associated with accessing third-party models through a cloud marketplace, making the AWS integration cost-neutral relative to direct API usage.
  • The three-tier GPT‑5.6 family (Luna/Terra/Sol) on Bedrock provides a structured cost-performance ladder within a single provider family, giving architects clear upgrade/downgrade paths without switching ecosystems.

When to Prefer It

  • Choose GPT‑5.6 Luna for high-volume, latency-sensitive pipelines such as real-time content tagging, bulk email classification, customer-service chatbot responses, or any workload where throughput and cost per call are the primary constraints.
  • Choose GPT‑5.6 Luna when building multi-step agentic workflows that invoke tools repeatedly, where the 80% price cut makes iterative tool-calling loops economically feasible at scale.
  • Choose GPT‑5.6 Terra for production applications requiring nuanced reasoning—such as summarization of complex documents, code review assistance, or structured data extraction—where Luna's capability ceiling is insufficient but Sol's cost is unjustifiable.
  • Choose GPT‑5.6 Sol (unchanged pricing) when the task demands frontier-level reasoning, advanced coding, cybersecurity analysis, or scientific research where output quality is the overriding concern and cost is secondary.
  • Prefer this Bedrock-hosted option over direct OpenAI API access when your organization requires AWS-native security controls, consolidated billing, audit logging via CloudTrail, or compliance with data residency requirements within supported AWS regions.
  • Consider these models for cost-optimization refactoring of existing workloads currently running on more expensive models (e.g., GPT‑5.5 or Sol), where the new Luna/Terra pricing may deliver acceptable quality at a fraction of the cost.

Availability

  • Status: Generally Available (GA) as of July 30, 2026; no preview or waitlist indicated.
  • Supported Regions: US East (N. Virginia), US East (Ohio), and US West (Oregon) only; no European or Asia-Pacific regions announced at this time.
  • API Endpoint: Accessible via the OpenAI Responses API on the bedrock-mantle endpoint within Amazon Bedrock.
  • Pricing Model: On-demand inference; GPT‑5.6 Luna reduced by 80%, GPT‑5.6 Terra reduced by 20% from prior rates; exact per-token prices available on the Amazon Bedrock pricing page.
  • Automatic Application: No customer action required—new prices apply automatically to all existing and new API calls from July 30, 2026.
  • Limitations: GPT‑5.6 Sol pricing is unchanged; batch inference pricing for GPT‑5.6 models is not mentioned in this announcement; regional availability is currently limited to three US regions.

Tags

Servicesbedrock
Typepricing
Conceptsllminference
Use Casescost-optimization
Providersopenai
GeographyAMERICAS

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.