Amazon EC2 Capacity Blocks for ML is now available in AWS GovCloud (US) Regions
Government and regulated-industry customers can now reserve cutting-edge NVIDIA B200/B300 GPU clusters inside FedRAMP-compliant GovCloud regions.
View original announcement →What's New
Amazon EC2 Capacity Blocks for ML is now generally available in both AWS GovCloud (US-West) and AWS GovCloud (US-East) regions, extending reserved GPU compute access to government agencies and regulated-industry customers. The service allows organizations to pre-reserve GPU instances—including the latest NVIDIA B200 and B300-based P6 instances—for defined durations to support machine learning workloads. This expansion brings the same assured, high-performance GPU capacity reservation capabilities previously available in commercial regions into FedRAMP-compliant, ITAR-controlled cloud environments.
How It Works
- Advance Reservation: Customers can reserve GPU capacity up to eight weeks in advance, with block durations ranging from short bursts up to 6 months, paying only for the reserved time window.
- Cluster Sizing: Each Capacity Block supports between 1 and 64 instances, with a maximum of 256 instances aggregated across multiple Capacity Blocks per account.
- UltraCluster Placement: Reserved instances are automatically co-located within Amazon EC2 UltraClusters, providing low-latency, petabit-scale, non-blocking networking optimized for distributed ML workloads.
- UltraServer Support: Capacity Blocks can also be used to reserve Amazon EC2 UltraServers, which interconnect multiple instances via a high-bandwidth accelerator fabric for the most memory- and compute-intensive AI/ML tasks.
- Multi-Account Sharing: Capacity Blocks can be shared across multiple AWS accounts using AWS Resource Access Manager (RAM), enabling centralized procurement with distributed consumption across teams or workloads.
- Instance Types in GovCloud: P6-B200 instances are available in GovCloud (US-West); both P6-B200 and P6-B300 instances are available in GovCloud (US-East).
Why It's Important
- Compliance-Ready GPU Compute: Government agencies and regulated industries (defense, intelligence, healthcare, finance) can now run sensitive ML workloads on cutting-edge GPU hardware within FedRAMP High and ITAR-compliant environments without leaving the GovCloud boundary.
- Eliminates GPU Scarcity Risk: Assured capacity reservation removes the uncertainty of on-demand GPU availability, which is critical for time-sensitive government AI programs with fixed project timelines or contract deliverables.
- Access to Latest NVIDIA Hardware: Availability of P6-B200 and P6-B300 (Blackwell architecture) instances gives government customers access to state-of-the-art accelerators for frontier model training and inference, closing the capability gap with commercial counterparts.
- Cost Efficiency for Bursty Workloads: The pay-for-duration model avoids the cost of always-on Reserved Instances while still providing predictability, making it economically viable for episodic training runs and prototype sprints.
- Organizational Coordination via RAM: Multi-account sharing enables large agencies or contractors to pool GPU investments across programs, maximizing utilization of reserved capacity across diverse workloads.
How It's Different
- vs. On-Demand Instances: On-demand provides no capacity guarantee and is subject to availability constraints; Capacity Blocks guarantee a specific cluster of GPU instances at a pre-scheduled time, eliminating launch failures during peak demand.
- vs. Reserved Instances (RIs): Standard RIs require long-term (1- or 3-year) commitments and charge continuously regardless of use; Capacity Blocks are purchased for exact durations (days to months) and only charge for that window.
- vs. Spot Instances: Spot instances can be interrupted at any time and are unsuitable for long training runs; Capacity Blocks provide uninterrupted, guaranteed access for the full reserved duration.
- vs. Commercial Region Capacity Blocks: GovCloud deployments operate within AWS's isolated, compliance-certified infrastructure, making this the only path to reserved GPU clusters that satisfy U.S. government data sovereignty and regulatory requirements.
- UltraCluster Networking: Unlike standard EC2 placement groups, Capacity Blocks guarantee physical co-location within UltraClusters, delivering petabit-scale non-blocking interconnects purpose-built for collective ML communication patterns (e.g., AllReduce).
When to Prefer It
- Pre-Training and Fine-Tuning Runs: When you have a defined training job that requires uninterrupted GPU access for days or weeks and cannot tolerate Spot interruptions or on-demand availability failures.
- Classified or Controlled ML Workloads: When data classification requirements (ITAR, CUI, FedRAMP High) mandate that compute must remain within GovCloud boundaries, ruling out commercial region alternatives.
- Scheduled Prototype Sprints: When a team needs GPU capacity for a bounded experiment window (e.g., a two-week model evaluation sprint) and wants cost certainty without a long-term RI commitment.
- Inference Demand Surges: When you anticipate a predictable spike in inference demand (e.g., a planned system launch or demonstration event) and need guaranteed GPU headroom at a specific future time.
- Multi-Program GPU Pooling: When a large agency or systems integrator manages multiple ML programs across accounts and wants to centrally reserve a GPU cluster and share it across teams via RAM to maximize utilization.
- UltraServer Workloads: When running extremely large models that exceed single-instance memory capacity and require tightly coupled multi-instance execution via the high-bandwidth accelerator interconnect.
Availability
- Status: Generally Available (GA) as of June 12, 2026.
- Supported Regions: AWS GovCloud (US-West) and AWS GovCloud (US-East); this announcement is specific to GovCloud expansion (the service was previously available in select commercial regions).
- Instance Types: P6-B200 in GovCloud (US-West); P6-B200 and P6-B300 in GovCloud (US-East).
- Reservation Window: Capacity Blocks can be reserved up to 8 weeks in advance for durations up to 6 months.
- Cluster Size Limits: 1 to 64 instances per Capacity Block; up to 256 instances total across all Capacity Blocks per account.
- Pricing Model: Pay-per-duration; customers are charged for the reserved time window only, not for continuous uptime like standard Reserved Instances.
- Multi-Account Sharing: Supported via AWS Resource Access Manager (RAM) at no additional service charge for the sharing mechanism itself.
- Documentation: Available at https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/ec2-capacity-blocks.html