← Back to all announcements
★★★★★ 29/04/2026

OpenAI GPT OSS and NVIDIA Nemotron Models Available on Amazon Bedrock in AWS GovCloud (US)

Government agencies and enterprises can now access cutting-edge open-weight models from OpenAI and NVIDIA in secure AWS GovCloud regions through a ...

View original announcement →

Visual Summary

graph TD A{{Bedrock in GovCloud - GPT OSS & Nemotron}}:::announced B(Amazon Bedrock):::compute C([Mantle Inference Engine]):::feature D([OpenAI API Compatibility]):::feature E([Automated Capacity Mgmt]):::feature F((Gov Agencies)):::external G((Enterprise Devs)):::external H([OpenAI GPT OSS Models]):::feature I([NVIDIA Nemotron Models]):::feature J(AWS GovCloud US):::compute F ==>|"compliant access"| A G ==>|"unified API"| A A -->|"hosted on"| B B -->|"runs within"| J A -->|"powered by"| C C -->|"enables"| D C -->|"provides"| E A -->|"serves"| H A -->|"serves"| I classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Amazon Bedrock now supports OpenAI's open-weight GPT OSS models (120B and 20B parameters) and NVIDIA's Nemotron model family (Nano 9B v2, Nano 12B v2, Nano 30B, and Super 120B) within the AWS GovCloud (US) regions. These additions expand the catalog of foundation models available through Bedrock's unified API, giving developers and enterprises access to high-performance open-weight models from two major AI providers. Notably, these models are powered by Mantle, a new distributed inference engine purpose-built for large-scale model serving on Amazon Bedrock.

How It Works

  • Both model families are served through Amazon Bedrock's standard serverless inference infrastructure, underpinned by the newly introduced Mantle distributed inference engine.
  • Mantle handles large-scale model serving by providing automated capacity management, unified resource pools, and sophisticated quality-of-service (QoS) controls that enable higher default customer quotas without manual intervention.
  • A key architectural feature is Mantle's out-of-the-box compatibility with the OpenAI API specification, meaning applications already written against the OpenAI API can invoke these models with minimal or no code changes.
  • Developers access the models through Bedrock's single, unified API, allowing seamless model switching across OpenAI GPT OSS, NVIDIA Nemotron, and other Bedrock-supported models without modifying application logic.

Why It's Important

  • For government agencies, defense contractors, and regulated enterprises operating under FedRAMP, ITAR, or other compliance frameworks, the availability of these models in AWS GovCloud (US) is significant because it allows them to leverage state-of-the-art open-weight foundation models without leaving the compliance boundary.
  • The open-weight nature of both GPT OSS and Nemotron models provides transparency into model architecture and weights, which is critical for organizations that require auditability and explainability.
  • NVIDIA Nemotron's range of SLM and LLM sizes also enables cost-efficient deployment of agentic AI workloads at varying compute budgets, while the Mantle engine's automated capacity management reduces operational overhead for teams scaling generative AI applications.

How It's Different

  • Previously, AWS GovCloud (US) had a more limited selection of foundation models on Bedrock compared to standard commercial regions, creating a capability gap for government and regulated-industry customers.
  • The introduction of Mantle as the underlying inference engine represents a meaningful infrastructure shift from prior model onboarding approaches — it standardizes and accelerates how new models are integrated into Bedrock, improves reliability through unified capacity pools, and natively supports OpenAI API compatibility, which was not a built-in feature of earlier Bedrock infrastructure.
  • Compared to self-hosting these open-weight models on EC2 or SageMaker, Bedrock's serverless delivery eliminates the need to manage GPU infrastructure, patching, and scaling logic while still providing access to the same open weights.

When to Prefer It

  • Choose OpenAI GPT OSS models (120B or 20B) on Bedrock GovCloud when you need strong general-purpose language understanding and generation with the transparency of open weights, particularly if your team already has tooling or prompts built around OpenAI API conventions and wants to avoid refactoring.
  • Opt for NVIDIA Nemotron models when building specialized agentic AI systems that require a range of model sizes to balance latency, cost, and accuracy — for example, using Nano 9B or 12B v2 for high-throughput, low-latency agent subtasks and Super 120B for complex reasoning steps.
  • Both families are especially well-suited for GovCloud workloads where data sovereignty, compliance, and auditability are non-negotiable, and where the fully open weights, datasets, and training recipes of Nemotron provide the documentation trail required for risk assessments and authority-to-operate (ATO) processes.

Availability

  • These models are generally available (GA) in AWS GovCloud (US) regions as of April 29, 2026.
  • Supported regions are specifically the AWS GovCloud (US) partition, which includes GovCloud (US-East) and GovCloud (US-West), making them accessible to customers with workloads subject to U.S. government compliance requirements.
  • Access is provided through Amazon Bedrock's serverless on-demand inference model, with automated capacity management via Mantle reducing quota friction.
  • No specific mention of cross-region inference support or fine-tuning availability for these models was included in the announcement, so customers should verify those capabilities through the Bedrock model catalog and service documentation before designing workflows that depend on them.

Tags

Servicesbedrock
Typenew-modelregion-expansionga-launch
Conceptsgenaillminferenceagentic-ai
Use Casesenterprisegovernment
Providersopenainvidia
GeographyAMERICAS

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.