Announcing region expansion of G6 instances on SageMaker AI Inference
Government agencies can now run GenAI inference on faster NVIDIA L4 GPUs within GovCloud's compliance boundary, at up to 2x the performance of G4dn.
View original announcement →Visual Summary
What's New
AWS has expanded availability of Amazon EC2 G6 instances for SageMaker AI inference to the AWS GovCloud (US-East) region. G6 instances are powered by NVIDIA L4 Tensor Core GPUs and deliver up to 2x the deep learning inference performance compared to the previous-generation G4dn instances. This expansion enables government agencies and regulated organizations to run generative AI inference workloads—including language models, image generation, and computer vision—while satisfying strict compliance and data residency requirements.
How It Works
- G6 instances are equipped with up to 8 NVIDIA L4 Tensor Core GPUs, each providing 24 GB of GPU memory (up to 192 GB total GPU memory per instance), enabling large batch inference and multi-model hosting.
- The NVIDIA L4 GPUs feature fourth-generation Tensor Cores for accelerated matrix math and third-generation RT Cores for graphics workloads, making them efficient for both ML inference and rendering tasks.
- Instances are backed by third-generation AMD EPYC processors, supporting up to 192 vCPUs, up to 100 Gbps network bandwidth, and up to 7.52 TB of local NVMe SSD storage for high-throughput data pipelines.
- G6 instances are built on the AWS Nitro System, which provisions GPUs in pass-through mode, delivering near bare-metal GPU performance with strong security isolation.
- Fractionalized GPU sizes (G6f) are available with as little as 1/8 of an L4 GPU (3 GB GPU memory), allowing cost-optimized deployments for smaller models that don't require a full GPU.
- Inference endpoints are deployed through the standard SageMaker AI real-time inference, asynchronous inference, or batch transform APIs, with on-demand and SageMaker Savings Plans pricing options available.
Why It's Important
- Government agencies and contractors operating under FedRAMP, ITAR, or other federal compliance frameworks can now leverage modern GPU-accelerated generative AI inference without moving workloads outside of the GovCloud boundary.
- The 2x performance improvement over G4dn instances means agencies can serve more requests per second at the same cost, or reduce infrastructure spend for equivalent throughput.
- Data residency requirements are fully satisfied since all inference compute remains within the AWS GovCloud (US-East) region, which is physically and logically isolated from standard AWS regions.
- Expanding access to small-to-medium language models, image generation, and computer vision in GovCloud accelerates AI adoption in defense, intelligence, healthcare, and civilian government use cases.
- The availability of fractionalized GPU sizes reduces the barrier to entry for cost-sensitive government programs that need GPU acceleration but cannot justify a full L4 GPU per endpoint.
How It's Different
- G6 instances deliver up to 2x higher deep learning inference performance compared to G4dn instances (which use NVIDIA T4 GPUs), making them the preferred upgrade path for existing GovCloud inference workloads.
- The NVIDIA L4's fourth-generation Tensor Cores provide significantly improved INT8 and FP8 throughput compared to the T4's third-generation Tensor Cores, directly benefiting quantized model inference.
- Unlike G4dn, G6 instances offer fractionalized GPU profiles (G6f), enabling right-sizing at the GPU level rather than the instance level, which is a cost optimization capability not previously available in GovCloud.
- G6 instances support up to 100 Gbps network bandwidth versus 25 Gbps on G4dn, reducing data transfer bottlenecks for large model artifacts and high-throughput inference pipelines.
- The combination of AMD EPYC CPUs and NVMe local storage on G6 provides better CPU-to-GPU data feeding performance compared to the Intel Cascade Lake CPUs on G4dn instances.
When to Prefer It
- Choose G6 instances when deploying small-to-medium language models (e.g., 7B–13B parameter models) in GovCloud that fit within 24 GB of GPU memory per GPU and require low-latency real-time inference.
- Use G6 instances for image generation workloads (e.g., diffusion models) in GovCloud environments where compliance mandates data must not leave the GovCloud boundary.
- Prefer G6 over G4dn when migrating existing GovCloud inference endpoints to improve throughput or reduce per-inference cost without changing the SageMaker deployment model.
- Select G6f fractionalized sizes when running multiple small models simultaneously on a single instance to maximize GPU utilization and minimize cost for lighter inference workloads.
- Use G6 instances for computer vision inference pipelines (object detection, image classification, video analysis) in government programs that require FedRAMP High or ITAR compliance.
- Consider G6 when your workload leverages NVIDIA-optimized libraries such as TensorRT, CUDA, or cuDNN, as the L4 GPU's architecture is specifically tuned for these frameworks.
Availability
- Status: Generally Available (GA) — this is a region expansion of an existing GA instance type, not a preview.
- New Region: AWS GovCloud (US-East) — now supported for SageMaker AI inference endpoints.
- Previously Supported Regions: G6 instances were already available for SageMaker AI inference in multiple standard commercial AWS regions prior to this announcement.
- Pricing Model: On-demand pricing with no upfront commitment, and eligible for Amazon SageMaker Savings Plans for discounted rates in exchange for usage commitments; pricing details available at the SageMaker AI pricing page.
- GPU Memory Constraint: Workloads must fit within 24 GB of GPU memory per GPU; models requiring more than 192 GB total GPU memory (8× L4) are not supported on a single G6 instance.
- Instance Variants: Standard G6 (full and multi-GPU) and G6f (fractionalized GPU, as little as 3 GB GPU memory / 1/8 of an L4) are available depending on workload size.