← Back to all announcements
★★☆☆☆ 11/05/2026

Announcing Region Expansion of P6-B200 instances on SageMaker Studio notebooks

NVIDIA Blackwell B200 GPUs with 1,440 GB memory now available in SageMaker Studio notebooks—2x faster than P5en for LLM fine-tuning.

View original announcement →

Visual Summary

graph TD A{{P6-B200 on SageMaker Studio}}:::announced B(SageMaker Studio):::compute C([8x NVIDIA Blackwell GPUs]):::feature D([1440 GB HBM3e Memory]):::feature E(JupyterLab / CodeEditor):::compute F((ML Practitioners)):::external G(Amazon S3 / EFS):::storage H([LLM Fine-Tuning]):::feature I([Multi-Modal Reasoning]):::feature F ==>|"develops in"| E E ==>|"provisions"| A A -->|"managed by"| B A -->|"powered by"| C C -->|"provides"| D A -->|"enables"| H A -->|"supports"| I B -->|"integrates"| G H -.->|"reads data"| G classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

Amazon EC2 P6-B200 instances are now generally available in SageMaker Studio notebooks in the US East (N. Virginia) region, expanding access to NVIDIA Blackwell GPU-powered compute for interactive ML development. These instances feature 8 NVIDIA B200 GPUs with 1,440 GB of high-bandwidth GPU memory paired with 5th Generation Intel Xeon (Emerald Rapids) processors. Customers can use them directly within JupyterLab or CodeEditor environments to develop, fine-tune, and experiment with large foundation models.

How It Works

  • P6-B200 instances are provisioned as notebook kernels within SageMaker Studio, accessible through the JupyterLab or CodeEditor application interfaces without requiring separate EC2 cluster management.
  • Each instance is equipped with 8 NVIDIA Blackwell B200 GPUs, providing 1,440 GB of aggregate high-bandwidth memory (HBM3e), enabling in-memory hosting of very large model parameter sets during interactive sessions.
  • The underlying CPU is a 5th Generation Intel Xeon Scalable processor (Emerald Rapids), which handles data preprocessing, orchestration, and host-side compute tasks alongside the GPU workloads.
  • Users interact with models and training loops in real time through notebook cells, leveraging frameworks such as PyTorch, JAX, or TensorFlow with CUDA/NCCL backends optimized for Blackwell architecture.
  • SageMaker Studio manages instance lifecycle, IAM-based access control, and storage integration (e.g., Amazon EFS, S3), so researchers can focus on model experimentation rather than infrastructure provisioning.

Why It's Important

  • The 1,440 GB of GPU memory allows practitioners to load and fine-tune frontier-scale models—including 70B+ parameter LLMs and mixture-of-experts architectures—entirely within a single interactive notebook session, eliminating the need to shard across multiple nodes for experimentation.
  • Up to 2x training throughput improvement over P5en instances means iteration cycles are significantly shorter, accelerating the research-to-production pipeline for generative AI applications.
  • Native integration with SageMaker Studio lowers the operational barrier: teams without deep MLOps expertise can access state-of-the-art GPU hardware through a managed, familiar notebook interface.
  • Support for multi-modal reasoning models (text, image, video) broadens the range of enterprise use cases—such as copilots and content generation—that can be prototyped and validated interactively before committing to full training runs.

How It's Different

  • P6-B200 instances deliver up to 2x better AI training performance compared to P5en (NVIDIA H100) instances, reflecting the architectural leap from Hopper to Blackwell GPU microarchitecture.
  • The 1,440 GB aggregate GPU memory pool is substantially larger than the 640 GB available on an 8-GPU P5en instance, enabling larger batch sizes and eliminating gradient checkpointing trade-offs for many model sizes.
  • NVIDIA Blackwell introduces 5th-generation NVLink and NVSwitch interconnects, providing higher GPU-to-GPU bandwidth than the prior generation, which benefits multi-GPU tensor-parallel workloads common in LLM fine-tuning.
  • Unlike raw EC2 access, the SageMaker Studio integration provides built-in experiment tracking, managed kernels, and direct connectivity to AWS data services, reducing setup overhead compared to self-managed Blackwell instances.
  • The Emerald Rapids CPU pairing offers improved PCIe 5.0 bandwidth and higher core counts versus the Sapphire Rapids CPUs in P5en, reducing CPU-side bottlenecks in data-intensive preprocessing pipelines.

When to Prefer It

  • Choose P6-B200 Studio notebooks when interactively fine-tuning large foundation models (e.g., 70B–400B parameter LLMs or MoE models) where the full model must reside in GPU memory during experimentation.
  • Prefer these instances for rapid prototyping of multi-modal generative AI applications—such as video generation or vision-language models—that require high memory capacity and compute throughput within a single session.
  • Use P6-B200 when iteration speed is critical and the 2x performance gain over P5en meaningfully reduces the feedback loop between hypothesis and result during model development.
  • Ideal for enterprise teams building copilots or domain-specific generative AI tools who need a managed, secure notebook environment without provisioning and configuring raw GPU clusters.
  • Select this option when working with mixture-of-experts architectures that have large sparse parameter counts, as the expanded GPU memory accommodates expert routing and activation without offloading to CPU RAM.

Availability

  • Status: Generally Available (GA) as of May 11, 2026.
  • Supported Region: AWS US East (N. Virginia) — this announcement represents a region expansion, implying prior availability in at least one other region.
  • Access Method: Available directly within SageMaker Studio via JupyterLab and CodeEditor application environments.
  • Pricing: On-demand pricing applies; specific rates are listed on the SageMaker Studio pricing page (not disclosed in the announcement).
  • Setup: Configuration instructions are available in the SageMaker Studio developer guides for JupyterLab and CodeEditor applications.
  • Limitations: No additional regions beyond US East (N. Virginia) are confirmed in this announcement; availability in other regions has not been specified.

Tags

Servicessagemaker
Typeregion-expansionga-launch
Conceptstrainingfine-tuninggenaillmmultimodal
Use Casesdeveloper-toolsenterprise
Providersnvidia
GeographyAMERICAS

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.