← Back to all announcements
★★☆☆☆ 12/05/2026

Announcing Region Expansion of G6 instances on SageMaker Notebook Instances

NVIDIA L4 GPU notebooks with 2x faster inference are now available in 9 new Asia Pacific and European regions on SageMaker.

View original announcement →

Visual Summary

graph TD A{{G6 Instances on SageMaker Notebooks}}:::announced B(Amazon SageMaker Studio):::compute C(NVIDIA L4 GPUs):::compute D(AMD EPYC Processors):::compute E([GenAI Fine-tuning]):::feature F([DL Inference 2x Faster]):::feature G((Data Scientists)):::external H([9 New Regions]):::feature I([JupyterLab / CodeEditor]):::feature G ==>|"selects instance"| A A ==>|"runs within"| B A -->|"powered by"| C A -->|"hosted on"| D B -->|"launches"| I A -->|"enables"| E A -->|"delivers"| F A -.->|"expands to"| H classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Amazon EC2 G6 instances are now generally available on SageMaker notebook instances across nine additional regions in Asia Pacific and Europe. These instances are powered by NVIDIA L4 Tensor Core GPUs and offer twice the deep learning inference performance compared to the previous-generation G4dn instances. The expansion enables data scientists and ML engineers in these regions to interactively develop, fine-tune, and test generative AI and other ML workloads directly within SageMaker's notebook environment.

How It Works

  • G6 instances are backed by up to 8 NVIDIA L4 Tensor Core GPUs, each with 24 GB of GPU memory, providing up to 192 GB of total GPU memory per instance for memory-intensive model workloads.
  • The underlying CPU is a third-generation AMD EPYC processor, offering a high-performance host compute layer to complement the GPU resources.
  • Users select a G6 instance type when creating or updating a SageMaker notebook instance, and the environment is accessible via JupyterLab or CodeEditor applications within SageMaker Studio.
  • The instances support interactive workflows, meaning users can iteratively write code, run training loops, test inference endpoints, and inspect outputs in real time within the notebook interface.
  • G6 instances leverage the NVIDIA L4's INT8 and FP16 Tensor Core capabilities, which are specifically optimized for inference and fine-tuning tasks common in modern generative AI pipelines.

Why It's Important

  • Expanding to nine new regions reduces latency and data residency concerns for customers in Asia Pacific and Europe who previously had to route workloads to other regions to access G6 hardware.
  • The 2x inference performance improvement over G4dn translates directly into faster iteration cycles during interactive development, reducing the time between code changes and validated results.
  • Generative AI fine-tuning and inference are among the most GPU-memory-hungry workloads; the 24 GB per GPU on L4 makes it feasible to load larger models (e.g., 7B–13B parameter LLMs) without multi-instance workarounds.
  • Availability in SageMaker notebook instances lowers the barrier to entry—teams can prototype on the same GPU class they intend to use in production without managing raw EC2 infrastructure.
  • Broader regional availability supports enterprise compliance requirements, particularly for European customers subject to GDPR data locality obligations.

How It's Different

  • G6 instances use NVIDIA L4 GPUs rather than the NVIDIA T4 GPUs found in G4dn instances, delivering 2x better deep learning inference throughput at comparable or lower cost per inference operation.
  • The L4 GPU introduces fourth-generation Tensor Cores with improved sparsity support, enabling more efficient execution of pruned and quantized models compared to T4-based G4dn instances.
  • Each L4 GPU carries 24 GB of GDDR6 memory versus 16 GB on the T4, allowing larger models or larger batch sizes to fit entirely on a single GPU without offloading.
  • G6 pairs the NVIDIA L4 with AMD EPYC CPUs rather than the Intel Cascade Lake CPUs in G4dn, offering higher core counts and memory bandwidth on the host side.
  • Unlike G4dn, which was primarily positioned for inference at launch, G6 is explicitly supported for both interactive training (including generative AI fine-tuning) and inference within SageMaker notebooks.

When to Prefer It

  • Choose G6 when fine-tuning large language models (e.g., using LoRA or QLoRA techniques) interactively in a notebook, where the 24 GB per GPU enables loading 7B–13B parameter models on a single GPU.
  • Prefer G6 over G4dn when running real-time inference tests during development and iteration speed is critical, given the 2x inference performance advantage.
  • Use G6 for computer vision workloads involving large batch image processing or video frame inference, where the higher GPU memory and Tensor Core throughput reduce per-batch latency.
  • Select G6 for NLP and language translation prototyping when models exceed the 16 GB memory ceiling of G4dn's T4 GPU, avoiding the complexity of model sharding across multiple GPUs.
  • Opt for G6 in European or Asia Pacific regions when data residency regulations prohibit sending training data or model weights to other AWS regions where G6 was previously available.
  • Use G6 for recommender engine development when embedding tables and interaction matrices are large enough to benefit from the expanded GPU memory and faster matrix operations.

Availability

  • Status: Generally available (GA) as of May 12, 2026.
  • New Regions: Asia Pacific (Tokyo, Mumbai, Sydney) and Europe (London, Paris, Frankfurt, Stockholm, Zurich).
  • Service: Available on SageMaker notebook instances; accessible via JupyterLab and CodeEditor applications in SageMaker Studio.
  • Pricing: Billed at standard Amazon EC2 G6 on-demand instance rates plus SageMaker notebook instance overhead; specific pricing varies by region and instance size—consult the SageMaker pricing page for regional rates.
  • Limitations: No specific limitations were disclosed in the announcement; standard SageMaker notebook instance quotas and service limits apply per account and region.
  • Documentation: Setup instructions are available in the SageMaker developer guides for JupyterLab and CodeEditor applications.

Tags

Servicessagemaker
Typeregion-expansionga-launch
Conceptsfine-tuninginferencetraininggenainlpcomputer-visionrecommendation
Providersnvidia
GeographyAPJEMEA

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.