Amazon SageMaker Studio notebooks now support G7e instance types
Run LLMs and spatial AI workloads on NVIDIA Blackwell GPUs with 96 GB/GPU memory directly inside SageMaker Studio notebooks.
View original announcement →Visual Summary
What's New
Amazon SageMaker Studio notebooks now support G7e instance types, powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, enabling data scientists to run high-performance AI inference, spatial computing, and multi-GPU workloads directly within their interactive notebook environments. G7e instances bring up to 8 GPUs with 96 GB of memory each, 5th Generation Intel Xeon processors, and up to 1600 Gbps of EFA networking bandwidth to SageMaker Studio's JupyterLab and Code Editor applications. This availability is currently limited to select US regions.
How It Works
- G7e instances are selectable as the underlying compute when creating or switching instance types in SageMaker Studio JupyterLab spaces or Code Editor spaces, replacing or supplementing existing GPU instance options.
- Each G7e instance is backed by up to 8 NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, each with 96 GB of GPU memory, yielding up to 768 GB of total GPU memory per instance for large model loading.
- Instances support up to 192 vCPUs and up to 2 TiB of system memory, paired with 5th Generation Intel Xeon Scalable (Emerald Rapids) processors for CPU-side preprocessing and orchestration tasks.
- NVIDIA GPUDirect Peer-to-Peer (P2P) via PCIe enables direct GPU-to-GPU data transfers within a single instance, reducing CPU bottlenecks for multi-GPU inference and training workloads.
- For multi-node scenarios, G7e instances support NVIDIA GPUDirect RDMA with EFAv4 in EC2 UltraClusters, allowing low-latency memory access across nodes without CPU involvement.
- Up to 1600 Gbps of Elastic Fabric Adapter (EFA) networking bandwidth supports high-throughput data movement between nodes, critical for distributed model serving and training.
- Local NVMe SSD storage of up to 15.2 TB is available for fast dataset access and model checkpoint storage during notebook-based experimentation.
- SageMaker Studio notebooks run each space on a single EC2 instance backed by an EBS volume, so users can switch to a G7e instance type to immediately access Blackwell GPU capabilities without leaving the Studio environment.
Why It's Important
- Data scientists can now interactively develop, prototype, and test LLM inference, agentic AI, and multimodal generative AI workloads on cutting-edge Blackwell GPU hardware directly within familiar notebook interfaces, eliminating the need to provision separate EC2 instances.
- The 96 GB per GPU memory capacity allows researchers to load very large models (e.g., 70B+ parameter LLMs) across multiple GPUs within a single instance, enabling rapid experimentation without complex distributed setup.
- Spatial computing and graphics-plus-AI workloads—such as robotic simulation, digital twins, and avatar-based applications—gain access to the highest-performance GPU option available in SageMaker Studio, accelerating development cycles for these emerging use cases.
- The integration of EFAv4 and GPUDirect RDMA support means that even small-scale multi-node experiments initiated from notebooks can benefit from low-latency inter-node communication, bridging the gap between notebook prototyping and production cluster workloads.
- Up to 2.3x inference performance improvement over G6e instances means faster iteration loops for model evaluation and benchmarking directly in notebooks, reducing time-to-insight.
How It's Different
- G7e delivers 2x the GPU memory per GPU (96 GB vs. ~48 GB on G6e), allowing significantly larger models to be loaded without tensor parallelism across more nodes.
- GPU memory bandwidth is 1.85x higher than G6e, directly improving throughput for memory-bandwidth-bound inference workloads such as autoregressive LLM decoding.
- Inter-GPU communication bandwidth is up to 4x greater than G6e, making multi-GPU tensor parallelism substantially faster and more efficient within a single instance.
- EFA networking bandwidth is up to 4x higher than G6e (up to 1600 Gbps vs. ~400 Gbps), enabling much faster multi-node communication for distributed workloads.
- The NVIDIA RTX PRO 6000 Blackwell architecture introduces fourth-generation ray tracing cores and neural shader-optimized streaming processors, providing 1.7x RT core TFLOPs over G6e for spatial computing workloads—a capability not present in prior SageMaker Studio GPU options.
- CPU-to-GPU bandwidth is up to 4x higher than G6e, improving performance for recommender systems and RAG pipelines that require frequent host-to-device data transfers.
- G7e offers up to 1.27x the raw compute TFLOPs compared to G6e, providing a meaningful uplift even for compute-bound workloads.
When to Prefer It
- Choose G7e when prototyping or fine-tuning very large language models (e.g., 70B+ parameters) that require more than 48 GB of GPU memory per GPU to fit without aggressive quantization.
- Prefer G7e for agentic AI and multimodal generative AI inference workloads where real-time latency and high memory bandwidth are critical to meeting service-level requirements during development.
- Select G7e when developing spatial computing applications—such as digital twins, robotic simulations, or neural rendering pipelines—that combine graphics rendering and AI inference in the same workload.
- Use G7e for RAG and recommender system prototyping where high CPU-to-GPU bandwidth reduces the bottleneck of moving large embedding batches from host memory to GPU.
- Opt for G7e when running small-scale multi-node experiments from notebooks that need low-latency inter-node communication via GPUDirect RDMA, bridging notebook-based prototyping with production cluster behavior.
- G7e is the right choice when benchmarking inference performance on Blackwell-generation hardware before committing to a production deployment architecture on the same GPU family.
- Consider G7e for physical AI model development (e.g., robotics, simulation) where the combination of high GPU memory, ray tracing performance, and AI compute is required in a single interactive environment.
Availability
- Status: Generally Available (GA) as of June 23, 2026.
- Supported Regions: US East (N. Virginia), US East (Ohio), and US West (Oregon); not yet available in other AWS regions.
- Supported Applications: Available in SageMaker Studio JupyterLab spaces and Code Editor spaces; not limited to Studio Classic.
- Pricing Model: On-demand pricing based on instance type and duration of use, consistent with standard SageMaker Studio notebook instance pricing; no upfront commitment required. Savings Plans may apply.
- Free Tier: The SageMaker AI Free Tier covers only ml.t3.medium instances and does not apply to G7e instances.
- Limitations: Multi-node GPUDirect RDMA with EFAv4 requires EC2 UltraClusters, which is a separate infrastructure consideration beyond standard Studio notebook single-instance usage.
Related Resources
- https://aws.amazon.com/ec2/instance-types/g7e/
- https://docs.aws.amazon.com/sagemaker/latest/dg/studio-updated-jl.html
- https://docs.aws.amazon.com/sagemaker/latest/dg/code-editor.html
- https://docs.aws.amazon.com/sagemaker/latest/dg/studio-updated.html
- https://aws.amazon.com/sagemaker/ai/pricing/?refid=ft_sagemaker