← Back to all announcements
★★★★☆ 30/07/2026

Gemma 4 models are now available on Amazon Bedrock in AWS GovCloud (US-West)

Government workloads can now use Gemma 4's reasoning, multimodal, and agentic capabilities inside FedRAMP-compliant AWS GovCloud.

View original announcement →

Visual Summary

graph TD A{{Gemma 4 on Bedrock GovCloud}}:::announced B(Amazon Bedrock):::compute C([Gemma 4 31B - Dense]):::feature D([Gemma 4 26B-A4B - MoE]):::feature E([Gemma 4 E2B - Compact]):::feature F([Native Function Calling]):::feature G([Multimodal Input]):::feature H((Gov/Regulated Users)):::external I(AWS GovCloud US-West):::compute H ==>|"API requests"| B B ==>|"hosts"| A A -->|"reasoning/coding"| C A -->|"cost-efficient"| D A -->|"low-latency"| E A -->|"tool use"| F A -.->|"text/image/video"| G I -->|"FedRAMP compliant"| B classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Google DeepMind's Gemma 4 family of open-weight models is now available on Amazon Bedrock in AWS GovCloud (US-West), bringing advanced generative AI capabilities to regulated and government workloads. The release includes three model variants — Gemma 4 31B, Gemma 4 26B-A4B, and Gemma 4 E2B — spanning dense and mixture-of-experts (MoE) architectures. These models support multimodal input, built-in reasoning, native function calling, and 35+ languages, running on a new Bedrock infrastructure innovation optimized for price-performance.

How It Works

  • Gemma 4 31B is a 30.7-billion parameter dense model with a 256K-token context window, optimized for reasoning- and coding-heavy workloads with built-in reasoning and native function calling.
  • Gemma 4 26B-A4B is a mixture-of-experts (MoE) model with 25.2B total parameters but only 3.8B active per token, enabling cost-efficient inference while retaining strong capability for latency-sensitive workloads.
  • Gemma 4 E2B is a compact model with 5.1B total parameters and 2.3B effective parameters, designed for low-latency interactive applications with a 128K-token context window.
  • All three variants support multimodal input across text and image (with the announcement noting broader support for video and audio), enabling diverse application types beyond text-only workflows.
  • Models run on a new Amazon Bedrock infrastructure innovation purpose-built for price-performance, with enhanced support for tool calling, structured output, reasoning, and response streaming.
  • Native function calling and structured output support allow developers to build reliable agentic pipelines and software engineering workflows directly on top of these open-weight models.
  • Access is managed through the standard Amazon Bedrock API and console, with model detail pages available in AWS documentation for configuration and integration guidance.

Why It's Important

  • GovCloud availability means federal agencies, defense contractors, and regulated industries (healthcare, finance) can now leverage state-of-the-art open-weight models within a FedRAMP-authorized, ITAR-compliant environment.
  • Open-weight models in a managed service give organizations the flexibility of open-source AI without the operational burden of self-hosting, combining customizability with AWS's enterprise security and scalability.
  • Built-in reasoning and native function calling reduce the engineering overhead required to build agentic and multi-step AI workflows, accelerating time-to-production for complex applications.
  • Multimodal support (text, image, and broader media types) expands the range of government and enterprise use cases addressable within a single, compliant environment.
  • MoE architecture availability (26B-A4B) provides a cost-efficient path to high-quality inference, making advanced AI economically viable for high-volume or budget-constrained government programs.
  • The 256K-token context window on Gemma 4 31B enables processing of long documents — such as legal filings, technical manuals, or policy documents — in a single inference call.

How It's Different

  • GovCloud-first availability distinguishes this from most commercial AI model launches, which typically reach GovCloud regions weeks or months after commercial regions, if at all.
  • MoE architecture (26B-A4B) offers a unique efficiency profile compared to purely dense models: near-full-model quality at a fraction of the active-parameter compute cost per token.
  • Open-weight licensing contrasts with proprietary models on Bedrock (e.g., Claude, Titan), giving organizations greater transparency, auditability, and potential for fine-tuning or offline deployment.
  • New Bedrock price-performance infrastructure is specifically called out as a platform innovation accompanying this launch, suggesting optimized hardware/software co-design beyond standard model hosting.
  • Three-tier model family (31B dense, 26B MoE, E2B compact) provides a single-vendor, single-API solution covering the full spectrum from high-accuracy to low-latency use cases, reducing integration complexity.
  • Compared to Gemma 3 models already on Bedrock (12B, 27B, 4B), Gemma 4 adds native reasoning, function calling, and significantly larger context windows as first-class features rather than prompt-engineered workarounds.

When to Prefer It

  • Federal and government agencies requiring FedRAMP High or ITAR-compliant AI inference should prefer Gemma 4 on Bedrock GovCloud over commercial-region alternatives.
  • Reasoning and code generation workloads (e.g., automated code review, security analysis, policy interpretation) benefit most from Gemma 4 31B's dense architecture and 256K context window.
  • Cost-sensitive, high-throughput applications such as document triage, classification pipelines, or real-time summarization are well-served by Gemma 4 26B-A4B's MoE efficiency.
  • Interactive, low-latency applications — chatbots, copilots, or real-time decision support tools — should target Gemma 4 E2B for its compact footprint and fast response times.
  • Agentic and tool-use workflows that require reliable structured output and function calling (e.g., automated workflows, API orchestration) benefit from the native support baked into all Gemma 4 variants.
  • Multilingual government or international programs supporting 35+ languages can leverage Gemma 4 without additional translation layers or separate model deployments.
  • Organizations evaluating open-weight vs. proprietary models can use Bedrock's model evaluation tools to benchmark Gemma 4 against Claude or other models on the same infrastructure before committing.

Availability

  • Status: Generally Available (GA) as of July 30, 2026.
  • Region: AWS GovCloud (US-West) only at launch; commercial region availability should be verified separately via the Amazon Bedrock model catalog.
  • Models available: Gemma 4 31B (dense, 256K context), Gemma 4 26B-A4B (MoE, 256K context), and Gemma 4 E2B (compact, 128K context).
  • Pricing model: On-demand inference pricing through Amazon Bedrock; specific per-token rates should be confirmed on the Amazon Bedrock pricing page, as GovCloud pricing may differ from commercial regions.
  • Multimodal scope: Documentation confirms text and image input for all three variants; the announcement references video and audio support, which should be verified against current model card details before relying on those modalities in production.
  • Access: Available via the Amazon Bedrock console and API; model detail pages are linked in the AWS documentation under Google model cards.
  • Prerequisites: Standard AWS GovCloud account with Amazon Bedrock access enabled; model access may require explicit enablement in the Bedrock console.

Tags

Servicesbedrock
Typenew-modelregion-expansion
Conceptsgenaillmmultimodalagentic-aicoding-assistantinference
Use Casesgovernmentopen-source
Providersgoogle
GeographyAMERICAS

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.