Gemma 4 models are now available on Amazon Bedrock in AWS GovCloud (US-West)
Government workloads can now use Gemma 4's reasoning, multimodal, and agentic capabilities inside FedRAMP-compliant AWS GovCloud.
View original announcement →Visual Summary
What's New
Google DeepMind's Gemma 4 family of open-weight models is now available on Amazon Bedrock in AWS GovCloud (US-West), bringing advanced generative AI capabilities to regulated and government workloads. The release includes three model variants — Gemma 4 31B, Gemma 4 26B-A4B, and Gemma 4 E2B — spanning dense and mixture-of-experts (MoE) architectures. These models support multimodal input, built-in reasoning, native function calling, and 35+ languages, running on a new Bedrock infrastructure innovation optimized for price-performance.
How It Works
- Gemma 4 31B is a 30.7-billion parameter dense model with a 256K-token context window, optimized for reasoning- and coding-heavy workloads with built-in reasoning and native function calling.
- Gemma 4 26B-A4B is a mixture-of-experts (MoE) model with 25.2B total parameters but only 3.8B active per token, enabling cost-efficient inference while retaining strong capability for latency-sensitive workloads.
- Gemma 4 E2B is a compact model with 5.1B total parameters and 2.3B effective parameters, designed for low-latency interactive applications with a 128K-token context window.
- All three variants support multimodal input across text and image (with the announcement noting broader support for video and audio), enabling diverse application types beyond text-only workflows.
- Models run on a new Amazon Bedrock infrastructure innovation purpose-built for price-performance, with enhanced support for tool calling, structured output, reasoning, and response streaming.
- Native function calling and structured output support allow developers to build reliable agentic pipelines and software engineering workflows directly on top of these open-weight models.
- Access is managed through the standard Amazon Bedrock API and console, with model detail pages available in AWS documentation for configuration and integration guidance.
Why It's Important
- GovCloud availability means federal agencies, defense contractors, and regulated industries (healthcare, finance) can now leverage state-of-the-art open-weight models within a FedRAMP-authorized, ITAR-compliant environment.
- Open-weight models in a managed service give organizations the flexibility of open-source AI without the operational burden of self-hosting, combining customizability with AWS's enterprise security and scalability.
- Built-in reasoning and native function calling reduce the engineering overhead required to build agentic and multi-step AI workflows, accelerating time-to-production for complex applications.
- Multimodal support (text, image, and broader media types) expands the range of government and enterprise use cases addressable within a single, compliant environment.
- MoE architecture availability (26B-A4B) provides a cost-efficient path to high-quality inference, making advanced AI economically viable for high-volume or budget-constrained government programs.
- The 256K-token context window on Gemma 4 31B enables processing of long documents — such as legal filings, technical manuals, or policy documents — in a single inference call.
How It's Different
- GovCloud-first availability distinguishes this from most commercial AI model launches, which typically reach GovCloud regions weeks or months after commercial regions, if at all.
- MoE architecture (26B-A4B) offers a unique efficiency profile compared to purely dense models: near-full-model quality at a fraction of the active-parameter compute cost per token.
- Open-weight licensing contrasts with proprietary models on Bedrock (e.g., Claude, Titan), giving organizations greater transparency, auditability, and potential for fine-tuning or offline deployment.
- New Bedrock price-performance infrastructure is specifically called out as a platform innovation accompanying this launch, suggesting optimized hardware/software co-design beyond standard model hosting.
- Three-tier model family (31B dense, 26B MoE, E2B compact) provides a single-vendor, single-API solution covering the full spectrum from high-accuracy to low-latency use cases, reducing integration complexity.
- Compared to Gemma 3 models already on Bedrock (12B, 27B, 4B), Gemma 4 adds native reasoning, function calling, and significantly larger context windows as first-class features rather than prompt-engineered workarounds.
When to Prefer It
- Federal and government agencies requiring FedRAMP High or ITAR-compliant AI inference should prefer Gemma 4 on Bedrock GovCloud over commercial-region alternatives.
- Reasoning and code generation workloads (e.g., automated code review, security analysis, policy interpretation) benefit most from Gemma 4 31B's dense architecture and 256K context window.
- Cost-sensitive, high-throughput applications such as document triage, classification pipelines, or real-time summarization are well-served by Gemma 4 26B-A4B's MoE efficiency.
- Interactive, low-latency applications — chatbots, copilots, or real-time decision support tools — should target Gemma 4 E2B for its compact footprint and fast response times.
- Agentic and tool-use workflows that require reliable structured output and function calling (e.g., automated workflows, API orchestration) benefit from the native support baked into all Gemma 4 variants.
- Multilingual government or international programs supporting 35+ languages can leverage Gemma 4 without additional translation layers or separate model deployments.
- Organizations evaluating open-weight vs. proprietary models can use Bedrock's model evaluation tools to benchmark Gemma 4 against Claude or other models on the same infrastructure before committing.
Availability
- Status: Generally Available (GA) as of July 30, 2026.
- Region: AWS GovCloud (US-West) only at launch; commercial region availability should be verified separately via the Amazon Bedrock model catalog.
- Models available: Gemma 4 31B (dense, 256K context), Gemma 4 26B-A4B (MoE, 256K context), and Gemma 4 E2B (compact, 128K context).
- Pricing model: On-demand inference pricing through Amazon Bedrock; specific per-token rates should be confirmed on the Amazon Bedrock pricing page, as GovCloud pricing may differ from commercial regions.
- Multimodal scope: Documentation confirms text and image input for all three variants; the announcement references video and audio support, which should be verified against current model card details before relying on those modalities in production.
- Access: Available via the Amazon Bedrock console and API; model detail pages are linked in the AWS documentation under Google model cards.
- Prerequisites: Standard AWS GovCloud account with Amazon Bedrock access enabled; model access may require explicit enablement in the Bedrock console.