← Back to all announcements
★★☆☆☆ 11/05/2026

Announcing Region Expansion of G6e instances on SageMaker Studio notebooks

G6e's 48 GB-per-GPU power is now available in 6 new regions, letting you fine-tune 13B LLMs locally without data residency compromises.

View original announcement →

Visual Summary

graph TD A{{G6e Instances on SageMaker Studio}}:::announced B(SageMaker Studio):::compute C(NVIDIA L40S GPUs):::compute D(AMD EPYC Processors):::compute E([JupyterLab / CodeEditor]):::feature F([GenAI Fine-tuning]):::feature G([LLM Deployment up to 13B]):::feature H([Diffusion Model Inference]):::feature I((Data Scientists)):::external I ==>|"launches"| E E ==>|"selects instance"| A A -->|"runs on"| B A -->|"powered by"| C A -->|"CPU compute"| D A -->|"enables"| F A -->|"supports"| G A -.->|"generates media"| H classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Amazon EC2 G6e instances are now generally available on SageMaker Studio notebooks in six additional regions: Middle East (Dubai), Asia Pacific (Tokyo, Seoul), and Europe (Frankfurt, Stockholm, Spain). These instances feature up to 8 NVIDIA L40S Tensor Core GPUs with 48 GB of GPU memory each, backed by third-generation AMD EPYC processors. The expansion enables developers in these regions to interactively fine-tune generative AI models, deploy LLMs up to 13B parameters, and generate images, video, and audio using diffusion models directly within SageMaker Studio's JupyterLab and CodeEditor environments.

How It Works

  • G6e instances are powered by up to 8 NVIDIA L40S Tensor Core GPUs, each with 48 GB of dedicated GPU memory, providing up to 384 GB of total GPU memory per instance for memory-intensive workloads.
  • Third-generation AMD EPYC processors handle CPU-side computation, complementing the GPU workloads for data preprocessing, orchestration, and mixed compute tasks.
  • Users select G6e instance types when launching JupyterLab or CodeEditor applications within SageMaker Studio, enabling interactive notebook-based workflows without separate cluster provisioning.
  • The instances support interactive model training workflows, allowing data scientists to iteratively fine-tune generative AI models and observe results in real time within the notebook environment.
  • For inference testing, G6e instances can host LLMs with up to 13B parameters and run diffusion model inference for multimodal content generation (images, video, audio) directly from the notebook.

Why It's Important

  • Developers in major financial, enterprise, and research hubs — Tokyo, Seoul, Frankfurt, Stockholm, Madrid, and Dubai — can now access high-performance GPU compute without routing workloads to distant regions, reducing latency and addressing data residency requirements.
  • The 2.5x performance improvement over G5 instances means faster iteration cycles for generative AI fine-tuning, directly reducing the time-to-insight for ML practitioners.
  • Having GPU-backed interactive notebooks lowers the barrier to entry for generative AI experimentation, allowing teams to prototype, fine-tune, and test LLMs and diffusion models in a single, managed environment.
  • Regional availability supports compliance and sovereignty requirements for organizations in regulated industries (finance, healthcare, government) in Europe and the Middle East that cannot move data across geographic boundaries.
  • The ability to interactively test model deployment on the same instance type used in production reduces the risk of performance surprises when moving from experimentation to deployment.

How It's Different

  • G6e instances deliver up to 2.5x better performance than the previous-generation G5 instances (which use NVIDIA A10G GPUs), making them significantly more capable for large-scale generative AI workloads.
  • The NVIDIA L40S GPU offers 48 GB of memory per GPU compared to 24 GB on the A10G in G5 instances, enabling larger models and larger batch sizes without memory bottlenecks.
  • Unlike general-purpose GPU instances, G6e is specifically optimized for both training and inference of modern generative AI workloads, including transformer-based LLMs and diffusion models.
  • The integration into SageMaker Studio notebooks differentiates this from raw EC2 access by providing a fully managed, IDE-like experience with built-in kernel management, experiment tracking, and AWS service integrations.
  • The combination of AMD EPYC CPUs and NVIDIA L40S GPUs offers a heterogeneous compute profile that balances cost-efficiency with raw GPU throughput compared to purely NVIDIA-CPU-paired alternatives.

When to Prefer It

  • Choose G6e on SageMaker Studio when fine-tuning open-source LLMs (e.g., Llama, Mistral) with up to 13B parameters interactively, where rapid iteration and real-time feedback are critical.
  • Prefer G6e when working with diffusion models for image, video, or audio generation tasks that require large GPU memory buffers to hold model weights and intermediate activations.
  • Use G6e when your organization operates under data residency or compliance requirements that mandate compute remain within specific regions such as EU (Frankfurt, Stockholm, Spain), Middle East (Dubai), or APAC (Tokyo, Seoul).
  • Select G6e over G5 when workloads are GPU memory-bound and previously required model sharding or quantization workarounds due to the 24 GB per-GPU limit of G5 instances.
  • Opt for G6e in SageMaker Studio when you need to interactively validate model deployment behavior before committing to a full SageMaker Endpoint deployment, reducing wasted inference infrastructure costs.
  • Consider G6e for multimodal generative AI prototyping pipelines that combine text, image, and audio generation in a single notebook session requiring sustained high GPU throughput.

Availability

  • Status: Generally Available (GA) as of May 11, 2026.
  • New Regions: Middle East (Dubai), Asia Pacific (Tokyo), Asia Pacific (Seoul), Europe (Frankfurt), Europe (Stockholm), Europe (Spain).
  • Supported Environments: Available for JupyterLab and CodeEditor applications within Amazon SageMaker Studio notebooks.
  • Instance Specs: Up to 8 NVIDIA L40S Tensor Core GPUs, 48 GB GPU memory per GPU (up to 384 GB total), third-generation AMD EPYC processors.
  • Pricing: Region-specific pricing is available on the AWS SageMaker pricing page; costs vary by instance size and region.
  • Limitations: Support is scoped to SageMaker Studio notebook applications; availability for SageMaker Training Jobs or Inference Endpoints in these regions may differ and should be verified separately.

Tags

Servicessagemaker
Typeregion-expansionga-launch
Conceptstraininginferencefine-tuninggenaillm
Use Casesdeveloper-toolsmulti-region
Providersnvidia
GeographyAPJEMEA

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.