← Back to all announcements
★★☆☆☆ 23/07/2026

Announcing region expansion of G6 instances on SageMaker AI Inference

Government agencies can now run GenAI inference on faster NVIDIA L4 GPUs within GovCloud's compliance boundary, at up to 2x the performance of G4dn.

View original announcement →

Visual Summary

graph TD A{{G6 Instances on SageMaker AI}}:::announced B(Amazon SageMaker AI):::compute C(NVIDIA L4 GPUs):::compute D(AWS GovCloud US-East):::compute E([Real-Time Inference]):::feature F([Fractionalized GPU - G6f]):::feature G([2x Performance vs G4dn]):::feature H((Government Agencies)):::external I(AWS Nitro System):::compute H ==>|"deploy endpoints"| A A ==>|"runs on"| B A -->|"powered by"| C A -->|"available in"| D B -->|"serves via"| E A -.->|"cost-optimized"| F C -->|"delivers"| G A -->|"built on"| I classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

AWS has expanded availability of Amazon EC2 G6 instances for SageMaker AI inference to the AWS GovCloud (US-East) region. G6 instances are powered by NVIDIA L4 Tensor Core GPUs and deliver up to 2x the deep learning inference performance compared to the previous-generation G4dn instances. This expansion enables government agencies and regulated organizations to run generative AI inference workloads—including language models, image generation, and computer vision—while satisfying strict compliance and data residency requirements.

How It Works

  • G6 instances are equipped with up to 8 NVIDIA L4 Tensor Core GPUs, each providing 24 GB of GPU memory (up to 192 GB total GPU memory per instance), enabling large batch inference and multi-model hosting.
  • The NVIDIA L4 GPUs feature fourth-generation Tensor Cores for accelerated matrix math and third-generation RT Cores for graphics workloads, making them efficient for both ML inference and rendering tasks.
  • Instances are backed by third-generation AMD EPYC processors, supporting up to 192 vCPUs, up to 100 Gbps network bandwidth, and up to 7.52 TB of local NVMe SSD storage for high-throughput data pipelines.
  • G6 instances are built on the AWS Nitro System, which provisions GPUs in pass-through mode, delivering near bare-metal GPU performance with strong security isolation.
  • Fractionalized GPU sizes (G6f) are available with as little as 1/8 of an L4 GPU (3 GB GPU memory), allowing cost-optimized deployments for smaller models that don't require a full GPU.
  • Inference endpoints are deployed through the standard SageMaker AI real-time inference, asynchronous inference, or batch transform APIs, with on-demand and SageMaker Savings Plans pricing options available.

Why It's Important

  • Government agencies and contractors operating under FedRAMP, ITAR, or other federal compliance frameworks can now leverage modern GPU-accelerated generative AI inference without moving workloads outside of the GovCloud boundary.
  • The 2x performance improvement over G4dn instances means agencies can serve more requests per second at the same cost, or reduce infrastructure spend for equivalent throughput.
  • Data residency requirements are fully satisfied since all inference compute remains within the AWS GovCloud (US-East) region, which is physically and logically isolated from standard AWS regions.
  • Expanding access to small-to-medium language models, image generation, and computer vision in GovCloud accelerates AI adoption in defense, intelligence, healthcare, and civilian government use cases.
  • The availability of fractionalized GPU sizes reduces the barrier to entry for cost-sensitive government programs that need GPU acceleration but cannot justify a full L4 GPU per endpoint.

How It's Different

  • G6 instances deliver up to 2x higher deep learning inference performance compared to G4dn instances (which use NVIDIA T4 GPUs), making them the preferred upgrade path for existing GovCloud inference workloads.
  • The NVIDIA L4's fourth-generation Tensor Cores provide significantly improved INT8 and FP8 throughput compared to the T4's third-generation Tensor Cores, directly benefiting quantized model inference.
  • Unlike G4dn, G6 instances offer fractionalized GPU profiles (G6f), enabling right-sizing at the GPU level rather than the instance level, which is a cost optimization capability not previously available in GovCloud.
  • G6 instances support up to 100 Gbps network bandwidth versus 25 Gbps on G4dn, reducing data transfer bottlenecks for large model artifacts and high-throughput inference pipelines.
  • The combination of AMD EPYC CPUs and NVMe local storage on G6 provides better CPU-to-GPU data feeding performance compared to the Intel Cascade Lake CPUs on G4dn instances.

When to Prefer It

  • Choose G6 instances when deploying small-to-medium language models (e.g., 7B–13B parameter models) in GovCloud that fit within 24 GB of GPU memory per GPU and require low-latency real-time inference.
  • Use G6 instances for image generation workloads (e.g., diffusion models) in GovCloud environments where compliance mandates data must not leave the GovCloud boundary.
  • Prefer G6 over G4dn when migrating existing GovCloud inference endpoints to improve throughput or reduce per-inference cost without changing the SageMaker deployment model.
  • Select G6f fractionalized sizes when running multiple small models simultaneously on a single instance to maximize GPU utilization and minimize cost for lighter inference workloads.
  • Use G6 instances for computer vision inference pipelines (object detection, image classification, video analysis) in government programs that require FedRAMP High or ITAR compliance.
  • Consider G6 when your workload leverages NVIDIA-optimized libraries such as TensorRT, CUDA, or cuDNN, as the L4 GPU's architecture is specifically tuned for these frameworks.

Availability

  • Status: Generally Available (GA) — this is a region expansion of an existing GA instance type, not a preview.
  • New Region: AWS GovCloud (US-East) — now supported for SageMaker AI inference endpoints.
  • Previously Supported Regions: G6 instances were already available for SageMaker AI inference in multiple standard commercial AWS regions prior to this announcement.
  • Pricing Model: On-demand pricing with no upfront commitment, and eligible for Amazon SageMaker Savings Plans for discounted rates in exchange for usage commitments; pricing details available at the SageMaker AI pricing page.
  • GPU Memory Constraint: Workloads must fit within 24 GB of GPU memory per GPU; models requiring more than 192 GB total GPU memory (8× L4) are not supported on a single G6 instance.
  • Instance Variants: Standard G6 (full and multi-GPU) and G6f (fractionalized GPU, as little as 3 GB GPU memory / 1/8 of an L4) are available depending on workload size.

Tags

Servicessagemaker-ai
Typeregion-expansion
Conceptsinferencegenai
Use Casesgovernment
Providersnvidia
GeographyAMERICAS

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.