← Back to all announcements
★★☆☆☆ 11/05/2026

Amazon SageMaker Studio notebooks now support P5.4xl instance types

H100 GPUs are now available in SageMaker Studio notebooks, cutting training costs 40% and speeding up LLM and diffusion model work 4x.

View original announcement →

Visual Summary

graph TD A{{P5.4xl on SageMaker Studio}}:::announced B((Data Scientists)):::external C(SageMaker Studio):::compute D(NVIDIA H100 GPUs):::compute E([JupyterLab / CodeEditor]):::feature F([LLM & Diffusion Training]):::feature G([4x Faster / 40% Savings]):::feature H(Amazon EC2 P5.4xl):::compute I([GenAI Use Cases]):::feature B ==>|"launches notebook"| C C ==>|"selects instance"| A A -->|"provisions"| H H -->|"powered by"| D A -->|"runs in"| E E -->|"executes"| F F -->|"delivers"| G F -.->|"enables"| I classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Amazon EC2 P5.4xl instances, powered by NVIDIA H100 Tensor Core GPUs, are now generally available for use directly within Amazon SageMaker Studio notebooks. This brings one of AWS's most powerful GPU instance types into the interactive notebook environment, enabling data scientists and ML engineers to leverage H100 performance for deep learning and HPC workloads without leaving their development workflow. The availability spans multiple global regions, making high-performance generative AI experimentation and model training more accessible across a broader customer base.

How It Works

  • P5.4xl instances are selectable as the compute kernel when launching or switching instance types within SageMaker Studio's JupyterLab or CodeEditor applications, requiring no infrastructure provisioning outside the Studio interface.
  • Each P5.4xl instance is backed by NVIDIA H100 Tensor Core GPUs, which feature fourth-generation Tensor Cores and NVLink interconnects, delivering significantly higher FP8, FP16, and BF16 throughput compared to prior GPU generations.
  • SageMaker Studio manages the underlying EC2 instance lifecycle, allowing users to start, stop, and switch instances on demand, with compute costs accruing only when the instance is running.
  • Workloads run within the familiar Jupyter or CodeEditor notebook environment, meaning existing training scripts, HuggingFace pipelines, PyTorch/TensorFlow code, and diffusion model libraries can be executed directly without modification.
  • Users can access developer guides for JupyterLab and CodeEditor application setup to configure the correct kernel environments and IAM permissions needed to spin up P5.4xl instances within their Studio domain.

Why It's Important

  • Bringing H100-class GPUs into the interactive notebook experience dramatically shortens the iteration cycle for researchers experimenting with large language models (LLMs) and diffusion models, eliminating the need to submit batch training jobs just to test ideas.
  • The claimed 4x acceleration over previous-generation GPU instances (such as P4d/A100-based) means experiments that previously took hours can complete in a fraction of the time, directly improving researcher productivity.
  • A reported 40% reduction in ML training costs makes it economically viable to run more experiments and larger models within budget constraints, lowering the barrier to generative AI development.
  • Supporting use cases like question answering, code generation, image/video generation, and speech recognition in a single instance type simplifies infrastructure decisions for teams building multi-modal generative AI applications.
  • Availability in SageMaker Studio means teams benefit from integrated MLOps tooling—experiment tracking, model registry, and data access—alongside the raw compute power of H100 GPUs.

How It's Different

  • Unlike P4d instances (NVIDIA A100 GPUs), P5.4xl instances use NVIDIA H100 Tensor Core GPUs with fourth-generation Tensor Cores, delivering substantially higher throughput for transformer-based model training and inference.
  • The P5.4xl is a mid-tier variant of the P5 family, offering a balance between cost and performance compared to the full p5.48xlarge (8x H100), making H100 compute accessible for teams that don't need the full 8-GPU configuration.
  • Previous GPU instance availability in SageMaker Studio was largely limited to older generations (P3, P4); P5.4xl represents the first H100-based option natively available in the interactive Studio notebook environment.
  • Compared to running P5 instances via SageMaker Training Jobs, the Studio notebook integration provides an interactive, low-latency development experience rather than a batch-oriented workflow, enabling real-time debugging and iterative experimentation.
  • The H100's support for FP8 precision (in addition to BF16/FP16) enables more efficient training of large-scale models compared to A100-based instances, which lack native FP8 hardware support.

When to Prefer It

  • Choose P5.4xl in Studio notebooks when interactively developing, fine-tuning, or debugging large language models (e.g., Llama, Mistral, Falcon) where rapid iteration and real-time feedback are critical.
  • Prefer this instance type when training or running inference on diffusion models (e.g., Stable Diffusion, DALL-E style architectures) that require high GPU memory bandwidth and tensor compute throughput.
  • Use P5.4xl when your workload benefits from H100-specific features such as FP8 precision training or high-speed NVLink memory bandwidth, and a single GPU partition is sufficient.
  • This is the right choice for HPC workloads co-located with ML workflows, such as physics simulations or molecular dynamics, that require the raw floating-point performance of H100 hardware.
  • Opt for P5.4xl over submitting SageMaker Training Jobs when you need an interactive environment for exploratory research, hyperparameter sweeps with manual inspection, or step-through debugging of GPU kernels.
  • Consider P5.4xl when cost efficiency is a priority for H100 workloads—its 40% cost reduction claim over prior generations makes it preferable for sustained training runs that would otherwise be cost-prohibitive on older GPU types.

Availability

  • Status: Generally Available (GA) as of May 11, 2026.
  • Supported Regions: US East (N. Virginia), US East (Ohio), US West (Oregon), Asia Pacific (Mumbai), Asia Pacific (Tokyo), Asia Pacific (Jakarta), and South America (São Paulo).
  • Supported Environments: Available within SageMaker Studio's JupyterLab and CodeEditor applications; not explicitly listed for SageMaker Classic notebooks.
  • Pricing: On-demand pricing applies per the SageMaker Studio notebook instance pricing page; costs accrue only while the instance is active.
  • Limitations: Not yet available in all AWS regions (notably absent from EU regions and additional Asia Pacific regions such as Singapore and Sydney at launch).
  • Prerequisites: Users must have appropriate IAM permissions and SageMaker Studio domain configuration to select P5.4xl as a notebook instance type; service quota increases may be required for P5 instances in some accounts.

Tags

Servicessagemaker
Typega-launchregion-expansion
Conceptstraininginferencellmgenai
Providersnvidia
GeographyAMERICASAPJ

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.