← Back to all announcements
★★☆☆☆ 27/05/2026

Announcing Region Expansion of P4de instances on SageMaker Notebook Instances

P4de's 640GB GPU memory and 60% faster training are now available in Tokyo—cutting costs 20% vs. P4d for large-scale ML workloads.

View original announcement →

Visual Summary

graph TD A{{P4de on SageMaker Notebooks - Tokyo}}:::announced B(Amazon SageMaker):::compute C(EC2 P4de Instances):::compute D([8x NVIDIA A100 80GB GPUs]):::feature E([640GB GPU Memory]):::feature F(SageMaker JupyterLab):::compute G(SageMaker Code Editor):::compute H((ML Practitioners)):::external I([60% Better Training]):::feature J(Amazon S3):::storage H ==>|"launches"| A A ==>|"provisions"| C C -->|"powered by"| D D -->|"delivers"| E A -->|"runs on"| B B -->|"IDE access"| F B -->|"IDE access"| G E -->|"enables"| I A -.->|"stores artifacts"| J classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Amazon Web Services has expanded the availability of Amazon EC2 P4de instances to the Asia Pacific (Tokyo) region on SageMaker notebook instances, making them generally available. P4de instances are equipped with 8 NVIDIA A100 GPUs, each featuring 80GB of HBM2e GPU memory—double the memory of the existing P4d instances—for a total of 640GB of GPU memory per instance. This expansion enables customers in the Tokyo region to leverage significantly enhanced ML training performance directly within their SageMaker notebook workflows.

How It Works

  • P4de instances are backed by 8 NVIDIA A100 Tensor Core GPUs, each with 80GB of high-bandwidth HBM2e memory, delivering a total of 640GB of GPU memory per instance—twice the per-GPU memory of P4d instances.
  • SageMaker notebook instances provision these EC2 P4de instances as the underlying compute, running a Jupyter Notebook server pre-configured with the SageMaker Python SDK, Boto3, deep learning frameworks (PyTorch, TensorFlow), and other ML libraries.
  • Users can launch P4de-backed notebook instances through the SageMaker console, CLI, or SDK by selecting the appropriate instance type (e.g., ml.p4de.24xlarge) when creating or updating a notebook instance.
  • The increased GPU memory allows larger model parameters, bigger batch sizes, and higher-resolution datasets to reside entirely in GPU memory, reducing costly memory swapping and enabling more efficient distributed training.
  • SageMaker Studio's JupyterLab and Code Editor applications can also be configured to run on P4de instances, providing a full-featured IDE experience backed by this high-performance hardware.
  • Underlying networking leverages high-throughput, low-latency connectivity consistent with the P4 family, supporting distributed and multi-node training workloads at scale.

Why It's Important

  • The 60% improvement in ML training performance over P4d instances directly translates to faster model iteration cycles, enabling teams to experiment more rapidly and reduce time to market for AI-powered products.
  • A 20% reduction in cost-to-train compared to P4d means organizations can achieve better performance economics, lowering the total cost of ownership for large-scale training workloads.
  • Availability in Asia Pacific (Tokyo) is strategically significant for customers with data residency requirements, latency-sensitive workflows, or regulatory constraints that mandate compute resources remain within Japan or the broader APAC region.
  • The 640GB total GPU memory pool unlocks training of very large models—such as large language models (LLMs) and vision transformers—that previously required complex model parallelism strategies or could not fit on P4d instances at all.
  • Customers working with high-resolution imaging data (medical imaging, satellite imagery, industrial inspection) benefit directly from the expanded memory, enabling larger batch sizes of high-dimensional inputs without downsampling.

How It's Different

  • P4de provides 80GB HBM2e memory per GPU versus 40GB HBM2 per GPU on P4d—a 2× increase in per-GPU memory capacity that fundamentally changes what model sizes and dataset resolutions are feasible in a single training run.
  • Up to 60% better ML training performance compared to P4d, driven by the combination of higher memory bandwidth (HBM2e vs. HBM2) and larger memory capacity reducing data movement overhead.
  • 20% lower cost to train relative to P4d instances, meaning P4de delivers both superior performance and better price-performance—not a trade-off between the two.
  • Compared to P4d, P4de instances are better suited for memory-bound workloads where the bottleneck is GPU memory capacity rather than raw compute throughput alone.
  • Unlike P3/P3dn instances (the prior generation), the entire P4 family (including P4de) offers 400 Gbps instance networking and GPUDirect RDMA, enabling efficient multi-node scale-out; P4de extends this with the larger memory envelope.

When to Prefer It

  • Choose P4de when training large foundation models or LLMs where model parameters and optimizer states exceed the 40GB per-GPU memory limit of P4d instances, avoiding the need for aggressive model sharding.
  • Prefer P4de for computer vision workloads involving high-resolution images or video frames, where larger GPU memory allows bigger batch sizes without resolution downsampling that could degrade model accuracy.
  • Use P4de when operating under strict data residency or compliance requirements that mandate compute resources remain in the Asia Pacific (Tokyo) region and you need top-tier GPU training performance.
  • Select P4de for iterative research and experimentation workflows in SageMaker notebooks where faster training turnaround (up to 60% improvement) meaningfully accelerates the hypothesis-test-iterate cycle.
  • Opt for P4de when total training cost is a concern and you are currently using P4d—the 20% cost-to-train reduction means migrating to P4de can lower your ML infrastructure spend while simultaneously improving throughput.
  • Consider P4de for multi-modal training tasks (e.g., combining text, image, and audio) where the combined memory footprint of multiple modality encoders benefits from the larger 640GB aggregate GPU memory pool.

Availability

  • Status: Generally Available (GA) as of May 27, 2026.
  • Newly added region: Asia Pacific (Tokyo) — this announcement specifically expands P4de availability to this region on SageMaker notebook instances.
  • Prior regions: P4de instances were already available in other AWS regions before this expansion; Tokyo represents the latest addition.
  • Supported surfaces: Available on SageMaker notebook instances; also usable with SageMaker Studio's JupyterLab and Code Editor applications.
  • Pricing: Follows standard Amazon EC2 P4de on-demand and reserved instance pricing; SageMaker notebook instance pricing applies on top of the underlying EC2 cost — refer to the Amazon SageMaker Pricing page for current rates.
  • Instance type: ml.p4de.24xlarge is the expected SageMaker instance type designation for this hardware configuration.
  • Limitation: Availability may be subject to regional capacity; users should verify instance availability in the Tokyo region console before building production dependencies on this instance type.

Tags

Servicessagemaker
Typeregion-expansionga-launch
Conceptstraining
Providersnvidia
GeographyAPJ

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.