← Back to all announcements
★★★★☆ 13/07/2026

Gemma-4-E2B-it for is now available in Amazon SageMaker JumpStart

Google DeepMind's multimodal Gemma-4-E2B-it—with vision, audio, reasoning, and function calling—is now one click away on SageMaker JumpStart.

View original announcement →

Visual Summary

graph TD A{{Gemma-4-E2B-it on JumpStart}}:::announced B(SageMaker JumpStart):::compute C(SageMaker Studio):::compute D([Multimodal Processing]):::feature E([Built-in Reasoning]):::feature F([Native Function Calling]):::feature G([Code Generation]):::feature H((Developer)):::external I(SageMaker Python SDK):::compute H ==>|"deploys via"| C H -->|"programmatic access"| I C ==>|"one-click deploy"| B I -->|"deploys"| B B ==>|"hosts"| A A -->|"text/image/audio"| D A -->|"step-by-step"| E A -.->|"agentic workflows"| F A -.->|"generates code"| G classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Google DeepMind's Gemma-4-E2B-it, a multimodal instruction-tuned foundation model, is now available in Amazon SageMaker JumpStart as of July 13, 2026. The model supports text, image, and audio inputs with text output, and includes a built-in reasoning mode for step-by-step problem solving. Customers can deploy it with a few clicks via SageMaker Studio or the SageMaker Python SDK.

How It Works

  • Multimodal input processing: The model accepts text, image, and audio inputs simultaneously, enabling rich cross-modal understanding within a single inference call.
  • Built-in reasoning mode: A chain-of-thought reasoning mechanism allows the model to internally deliberate step-by-step before producing a final answer, improving accuracy on complex tasks.
  • Image understanding capabilities: Supports object detection, document parsing, screen/UI understanding, chart comprehension, and OCR out of the box without additional fine-tuning.
  • Video understanding: The model can process and reason over video content, extending its multimodal capabilities beyond static images.
  • Native function calling: Supports agentic workflows by enabling the model to invoke external tools or APIs directly, facilitating autonomous multi-step task execution.
  • Code capabilities: Handles code generation, completion, and correction across multiple programming languages.
  • Multilingual support: Operates across dozens of languages, making it suitable for globally distributed applications.
  • JumpStart deployment: Customers deploy via the SageMaker Studio Models landing page or the SageMaker Python SDK, with infrastructure provisioning handled automatically by JumpStart.

Why It's Important

  • Efficient edge-and-cloud flexibility: Optimized for efficient local execution, Gemma-4-E2B-it bridges the gap between on-device and cloud deployment, giving teams architectural flexibility without sacrificing capability.
  • Reduced time-to-deployment: JumpStart's one-click deployment eliminates the need to manually configure endpoints, container images, or instance types, dramatically lowering the barrier to production.
  • Broad task coverage in a single model: The combination of vision, audio, reasoning, code, and agentic capabilities means teams can address diverse use cases without managing multiple specialized models.
  • Agentic AI enablement: Native function calling support makes this model a strong candidate for building autonomous agents that interact with external systems, a rapidly growing enterprise need.
  • Google DeepMind quality on AWS infrastructure: Customers gain access to a frontier-class model from a leading AI lab while keeping data and workloads within their existing AWS security and compliance boundaries.

How It's Different

  • True multimodal input (text + image + audio): Unlike many models available in JumpStart that handle only text or text-and-image, Gemma-4-E2B-it adds audio as a native input modality.
  • Integrated reasoning mode: The built-in step-by-step reasoning capability is a first-class feature of the model architecture, not a prompt-engineering workaround, distinguishing it from standard instruction-tuned models.
  • Optimized for efficient execution: The "E2B" designation signals an efficiency-first design, making it more cost-effective to run at scale compared to larger parameter-count multimodal alternatives.
  • Native function calling: Many open-weight models require external scaffolding for tool use; Gemma-4-E2B-it supports function calling natively, simplifying agentic pipeline construction.
  • Video understanding included: Video comprehension is not commonly available in similarly sized open-weight models, giving this model a differentiated capability for media and surveillance use cases.

When to Prefer It

  • Document intelligence pipelines: When your application requires parsing invoices, forms, or scanned documents with OCR and layout understanding, this model handles it natively without a separate OCR service.
  • Agentic and tool-use workflows: When building AI agents that need to call APIs, query databases, or orchestrate multi-step tasks, the native function calling support simplifies the architecture significantly.
  • Cost-sensitive multimodal deployments: When you need multimodal capabilities but want to minimize inference costs, the efficiency-optimized design makes it preferable to larger frontier models.
  • Screen and UI automation: When building accessibility tools, RPA assistants, or UI testing agents that must interpret screenshots or application interfaces.
  • Multilingual enterprise applications: When your user base spans multiple languages and you need a single model to handle diverse linguistic inputs without separate localization models.
  • Reasoning-intensive tasks: When accuracy on multi-step logical, mathematical, or analytical questions is critical and you want the model to show its work before committing to an answer.
  • Video content analysis: When your use case involves summarizing, searching, or extracting insights from video content at scale within the AWS ecosystem.

Availability

  • GA status: Generally available as of July 13, 2026; no preview or beta restrictions mentioned.
  • Access method: Available through the SageMaker Studio Models landing page (SageMakerPublicHub) and the SageMaker Python SDK.
  • Regional availability: Specific supported regions are not listed in the announcement; customers should check the SageMaker JumpStart regional availability documentation for the latest region support.
  • Pricing model: Standard SageMaker endpoint pricing applies based on the instance type selected for deployment; no separate model licensing fee is indicated, consistent with open-weight model access in JumpStart.
  • License responsibility: As with all JumpStart third-party models, customers are responsible for reviewing and complying with the applicable Google DeepMind/Gemma model license terms before use.
  • Deployment prerequisite: Requires an active AWS account with appropriate SageMaker IAM permissions; access via SageMaker Studio or SDK.

Tags

Servicessagemaker-jumpstart
Typenew-modelga-launch
Conceptsmultimodalgenaiinferenceagentic-ai
Use Casesdeveloper-tools
Providersgoogle
GeographyGlobal

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.