Announcing region expansion of G7e instances on SageMaker AI inference
Deploy 70B-parameter LLMs on a single instance with 2.3x faster inference, now available in Seoul, Tokyo, and London.
View original announcement →Visual Summary
What's New
Amazon Web Services has expanded the availability of EC2 G7e instances for SageMaker AI inference to three new regions: Asia Pacific (Seoul), Asia Pacific (Tokyo), and Europe (London). These instances are powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs and deliver up to 2.3x inference performance compared to the previous-generation G6e instances. The expansion enables customers in Asia and Europe to deploy generative AI inference endpoints closer to their end users, reducing latency for production workloads.
How It Works
- G7e instances are equipped with up to 8 NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, each providing 96 GB of GPU memory, for a total of 768 GB of GPU memory per instance.
- The instances use 5th Generation Intel Xeon Scalable (Emerald Rapids) processors, supporting up to 192 vCPUs and up to 2 TiB of system memory.
- Up to 1,600 Gbps of Elastic Fabric Adapter (EFA) networking bandwidth is available, enabling high-throughput data transfer between instances and supporting distributed inference scenarios.
- G7e instances support up to 15.2 TB of local NVMe SSD storage, providing fast local I/O for model loading and intermediate data.
- NVIDIA GPUDirect P2P via PCIe enables high-bandwidth GPU-to-GPU communication (up to 4x compared to G6e), allowing large models to be efficiently sharded across multiple GPUs within a single instance.
- FP8 precision support allows models with up to 70B parameters to be served on a single G7e instance without requiring multi-node configurations, simplifying deployment architecture.
- SageMaker AI inference endpoints can be deployed on G7e instances using standard SageMaker real-time, asynchronous, or batch inference APIs with on-demand or Savings Plan pricing.
Why It's Important
- The regional expansion to Seoul, Tokyo, and London directly reduces inference latency for end users in Asia and Europe, which is critical for real-time generative AI applications such as chatbots, agentic AI, and multimodal services.
- The ability to host up to 70B parameter models on a single instance with FP8 precision eliminates the operational complexity and cost overhead of multi-node inference setups.
- With 768 GB of total GPU memory per instance, organizations can serve large foundation models without model sharding across nodes, improving reliability and reducing inter-node communication bottlenecks.
- The 2.3x inference performance improvement over G6e translates directly into higher throughput and lower cost-per-inference token, making large-scale LLM deployments more economically viable.
- Support for spatial computing, robotic simulation, digital twins, and scientific computing workloads broadens the addressable use cases beyond pure NLP, making G7e a versatile platform for next-generation AI applications.
How It's Different
- G7e provides 2x the GPU memory per GPU (96 GB vs. 48 GB) compared to G6e, enabling significantly larger models to fit on a single GPU or instance.
- GPU memory bandwidth is 1.85x higher than G6e, reducing memory bottlenecks during autoregressive token generation in LLM inference.
- Inter-GPU communication bandwidth is up to 4x greater than G6e, and EFA networking bandwidth is also 4x higher, supporting faster model parallelism and distributed inference.
- CPU-to-GPU bandwidth is up to 4x higher than G6e, which specifically benefits RAG pipelines and recommender systems that require frequent data transfers between host and accelerator memory.
- The fourth-generation NVIDIA ray tracing cores and neural shader-optimized streaming processors deliver 1.7x RT core TFLOPs compared to G6e, making G7e uniquely suited for graphics-plus-AI workloads that G6e cannot handle as efficiently.
- Unlike P-series instances optimized purely for training, G7e is purpose-built for cost-effective inference with a balance of GPU memory, bandwidth, and graphics capabilities not found in other SageMaker instance families.
When to Prefer It
- Choose G7e when deploying LLM inference for models in the 30B–70B parameter range that need to run on a single instance to avoid multi-node complexity and associated latency penalties.
- Prefer G7e for latency-sensitive generative AI applications serving users in Japan, South Korea, or the United Kingdom, where proximity to the new regions reduces round-trip time.
- Use G7e for multimodal or agentic AI workloads that combine text, image, and video generation and require both high GPU memory capacity and high memory bandwidth simultaneously.
- G7e is the right choice for spatial computing workloads such as digital twins, robotic simulation, and avatar-based assistants that require both ray tracing performance and AI inference on the same instance.
- Select G7e for RAG-based applications and recommender systems that benefit from the 4x improvement in CPU-to-GPU bandwidth, enabling faster retrieval and embedding lookups.
- Consider G7e for scientific computing workloads that require large GPU memory footprints and high-bandwidth inter-GPU communication within a single node.
- Prefer G7e over G6e whenever FP8 precision inference is required for large models, as the Blackwell architecture's native FP8 support maximizes throughput efficiency.
Availability
- Status: Generally Available (GA) for SageMaker AI inference endpoints.
- New Regions: Asia Pacific (Seoul), Asia Pacific (Tokyo), and Europe (London) as of July 23, 2026.
- Previously Supported Regions: Available in additional regions announced prior to this expansion (specific prior regions not enumerated in this announcement).
- Pricing Model: On-demand pricing with no upfront commitments, or Amazon SageMaker Savings Plans for committed usage discounts; specific per-hour rates available on the SageMaker AI pricing page.
- Supported Inference Modes: Compatible with SageMaker real-time inference, asynchronous inference, and batch transform endpoints.
- Instance Specifications: Available in configurations up to 8 GPUs (768 GB total GPU memory), 192 vCPUs, 2 TiB system memory, and 15.2 TB local NVMe SSD.
- Limitation: Multi-node G7e configurations for models exceeding 70B parameters (at FP8) would require additional architecture planning, as the single-instance limit is 768 GB GPU memory.