← Back to all announcements
★★★★☆ 13/07/2026

Voxtral-Mini-4B-Realtime for real-time speech transcription is now available in Amazon SageMaker JumpStart

Deploy Mistral's streaming speech model inside your own AWS account — real-time, multilingual, with tunable latency vs. accuracy trade-offs.

View original announcement →

Visual Summary

graph TD A{{Voxtral-Mini-4B-Realtime}}:::announced B(SageMaker JumpStart):::compute C(SageMaker Studio):::compute D(SageMaker Python SDK):::compute E([Real-Time Streaming]):::feature F([Multilingual 13 Languages]):::feature G([Configurable Latency]):::feature H(Amazon CloudWatch):::storage I((Developer)):::external I ==>|"deploys"| B B ==>|"hosts"| A C -->|"one-click deploy"| B D -->|"programmatic deploy"| B A -->|"transcribes"| E A -->|"supports"| F A -.->|"tunes"| G A -->|"logs to"| H classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

AWS has made Voxtral-Mini-4B-Realtime-2602, a multilingual real-time speech transcription model from Mistral AI, available in Amazon SageMaker JumpStart. The model features a natively streaming architecture designed for low-latency audio-to-text transcription across 13 languages. Customers can deploy it directly from SageMaker Studio or via the SageMaker Python SDK with minimal setup.

How It Works

  • Streaming-native architecture: The model is built from the ground up for real-time transcription, processing audio as a continuous stream rather than in discrete batches, enabling sub-second latency responses.
  • 4B parameter scale: At 4 billion parameters, the model balances transcription quality with computational efficiency, making it suitable for deployment on standard SageMaker inference instances.
  • Configurable transcription delay: Users can tune the trade-off between latency and accuracy by adjusting transcription delay parameters, allowing the model to emit partial transcripts sooner or wait for more audio context before committing.
  • Multilingual support: The model natively handles transcription across 13 languages without requiring separate language-specific models or explicit language-detection preprocessing.
  • SageMaker JumpStart deployment: The model is accessible via the SageMaker Studio Models landing page or programmatically through the SageMaker Python SDK, deploying to a managed real-time inference endpoint in the customer's AWS account.
  • Foundation model hosting: Once deployed, the endpoint follows standard SageMaker real-time inference patterns, supporting autoscaling, logging via CloudWatch, and integration with other AWS services.

Why It's Important

  • Low-latency speech applications become accessible: Developers can now build real-time captioning, live transcription, and voice-driven interfaces on AWS without managing complex streaming inference infrastructure from scratch.
  • Mistral AI's frontier audio model on AWS: Voxtral represents Mistral AI's entry into audio AI, and its availability on JumpStart brings a competitive, non-AWS-native model into the managed AWS ecosystem, expanding customer choice.
  • Reduces time-to-deployment: JumpStart's one-click deployment eliminates the need to containerize, configure, and optimize the model manually, cutting deployment time from days to minutes.
  • Configurable latency-accuracy trade-off: This is practically significant for use cases like live broadcasting (prioritize speed) versus medical transcription (prioritize accuracy), giving teams a single model that adapts to different SLA requirements.
  • Multilingual coverage in one endpoint: Supporting 13 languages from a single deployed model reduces infrastructure overhead for global applications compared to maintaining separate per-language models.

How It's Different

  • Natively streaming vs. batch-first models: Unlike models such as OpenAI Whisper (which processes fixed audio chunks) or Amazon Transcribe's standard mode, Voxtral-Mini-4B-Realtime is architecturally designed for continuous streaming, not retrofitted for it.
  • Configurable delay parameter: Most transcription services offer fixed latency profiles; Voxtral's explicit delay configuration gives engineers fine-grained control over the latency-accuracy curve at inference time.
  • Mistral AI provenance: This is a third-party frontier model (not an AWS-built model), distinguishing it from Amazon Transcribe and giving customers access to Mistral's research lineage and model characteristics.
  • Self-hosted on customer infrastructure: Unlike managed API services (Amazon Transcribe, Deepgram, AssemblyAI), the model runs inside the customer's AWS account, keeping audio data within their VPC and security boundary.
  • 4B parameter efficiency: Compared to larger speech models, the 4B parameter footprint targets a cost-performance sweet spot suitable for real-time workloads without requiring high-end GPU instances.

When to Prefer It

  • Live captioning and accessibility: When building real-time closed captioning for video conferencing, live events, or broadcast media where sub-second latency is a hard requirement.
  • Data residency and compliance requirements: When audio data cannot leave the customer's AWS account or VPC, making managed third-party transcription APIs unsuitable.
  • Multilingual contact center analytics: When processing customer calls across multiple languages in real time and a single unified model is preferable to routing to language-specific services.
  • Voice-driven agentic applications: When integrating speech input into LLM-based agents or real-time assistants where streaming transcription must feed downstream NLP pipelines with minimal delay.
  • Latency-accuracy tuning is needed: When different deployment environments (e.g., demo vs. production vs. medical) require different latency profiles from the same underlying model.
  • Cost-sensitive real-time workloads: When a 4B parameter model provides sufficient accuracy and the team wants to avoid the per-minute pricing of managed transcription APIs at high call volumes.

Availability

  • GA status: Generally available as of July 13, 2026, via Amazon SageMaker JumpStart.
  • Access method: Available through the SageMaker Studio Models landing page (SageMakerPublicHub) or the SageMaker Python SDK; no separate sign-up required beyond standard AWS account access.
  • Regional availability: Specific supported regions are not enumerated in the announcement; customers should check the SageMaker JumpStart model catalog within their target region for availability.
  • Pricing model: Standard SageMaker real-time inference pricing applies based on the instance type selected for the endpoint; there is no separate per-minute transcription charge, but Mistral AI license terms must be reviewed and accepted before use.
  • License compliance: As with all third-party JumpStart models, customers are responsible for reviewing and complying with Mistral AI's applicable license terms before deploying or using the model.
  • Existing endpoint continuity: Consistent with JumpStart policy, any endpoints deployed from this model will remain functional even if the model is later delisted from the catalog.

Tags

Servicessagemaker-jumpstart
Typenew-modelga-launch
Conceptsspeechnlpinferencemultimodal
Use Casesdeveloper-tools
Providersmistral
GeographyGlobal

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.