← Back to all announcements
★★★★☆ 18/06/2026

Ministral-3-14B-Instruct for multimodal reasoning and agentic AI is now available in Amazon SageMaker JumpStart

Mistral's compact 14B model brings multimodal vision, native function calling, and multilingual support to SageMaker JumpStart with one-click deployment.

View original announcement →

Visual Summary

graph TD A{{Ministral-3-14B on JumpStart}}:::announced B(SageMaker Studio):::compute C(SageMaker Python SDK):::compute D([Multimodal Inference]):::feature E([Agentic Function Calling]):::feature F([Multilingual Support]):::feature G((Developers)):::external H(SageMaker Endpoint):::compute I([Edge Deployment]):::feature G ==>|"deploys via"| B G -->|"programmatic"| C B ==>|"provisions"| A C -->|"provisions"| A A ==>|"hosts on"| H H -->|"serves"| D H -->|"enables"| E H -.->|"supports"| F A -.->|"optimized for"| I classDef announced fill:#ff9900,stroke:#ec7211,color:#fff,font-weight:bold classDef compute fill:#e3f2fd,stroke:#1565c0,color:#1565c0 classDef storage fill:#e8f5e9,stroke:#2e7d32,color:#2e7d32 classDef feature fill:#fff3e0,stroke:#e65100,color:#e65100 classDef external fill:#f5f5f5,stroke:#616161,color:#616161

What's New

AWS has made Ministral-3-14B-Instruct-2512 from Mistral AI available in Amazon SageMaker JumpStart, adding a compact yet capable multimodal foundation model to the platform's growing model catalog. The model combines vision understanding, native function calling, JSON output, and multilingual support within a 14B-parameter architecture explicitly optimized for edge deployment. Customers can deploy it immediately via SageMaker Studio's Models page or the SageMaker Python SDK with minimal configuration.

How It Works

  • Multimodal inference: The model accepts both image and text inputs, enabling it to analyze visual content and generate contextually grounded textual responses in a single inference call.
  • Agentic function calling: Native support for structured function calling and JSON output allows the model to act as an orchestrator or tool-using agent within multi-step workflows without additional prompt engineering scaffolding.
  • Multilingual processing: The model natively handles dozens of languages—including English, French, Spanish, German, Chinese, Japanese, Korean, and Arabic—making it suitable for globally distributed applications.
  • SageMaker JumpStart deployment: Customers deploy the model through the SageMaker Studio Models landing page (a few-click UI flow) or programmatically via the SageMaker Python SDK, which provisions a managed real-time inference endpoint backed by AWS infrastructure.
  • Incremental customization: As with other JumpStart models, customers can fine-tune the base model on domain-specific data before deployment, leveraging SageMaker's managed training infrastructure.
  • Edge-optimized architecture: The 14B-parameter size is deliberately compact relative to frontier-scale models, reducing memory footprint and inference latency to support deployment on resource-constrained or edge-adjacent hardware configurations.

Why It's Important

  • Lowers the barrier to multimodal AI: Packaging a vision-capable, agentic model inside JumpStart's one-click deployment experience removes the operational complexity of sourcing, containerizing, and serving a multimodal model from scratch.
  • Enables cost-efficient agentic pipelines: At 14B parameters, the model delivers function-calling and reasoning capabilities at a fraction of the compute cost of 70B+ models, making agentic architectures economically viable for a broader set of customers.
  • Expands Mistral AI's AWS footprint: This addition deepens the Mistral AI–AWS partnership and gives customers a managed, AWS-supported path to Mistral models beyond what is available through Amazon Bedrock.
  • Accelerates vision-enabled application development: Developers building document analysis, visual QA, or inspection automation use cases gain a production-ready endpoint without managing model weights, containers, or scaling logic.
  • Supports global product requirements: Built-in multilingual capability means teams can serve international users from a single model deployment rather than maintaining separate language-specific models.

How It's Different

  • Multimodal in a compact form factor: Unlike many small models (≤14B) that are text-only, Ministral-3-14B-Instruct combines image understanding with text generation, a capability typically associated with much larger models.
  • Native function calling vs. prompt-engineered workarounds: The model's built-in function-calling and JSON output support is a first-class feature, not a fine-tuned behavior, reducing reliability issues common in models that simulate tool use through prompting alone.
  • Edge-deployment optimization: The architecture is explicitly designed for lower-resource environments, distinguishing it from other multimodal models in JumpStart that target high-memory GPU instances.
  • Broader language coverage than typical small models: Support for Arabic, Chinese, Japanese, and Korean alongside European languages goes beyond the English-centric focus of many compact open-weight models.
  • Managed AWS deployment vs. self-hosted Mistral: Compared to running Mistral models on self-managed EC2 or EKS, JumpStart provides auto-scaling, endpoint monitoring, and IAM-integrated access control out of the box.

When to Prefer It

  • Edge or resource-constrained deployments: Choose this model when inference must run on smaller GPU instances (e.g., single A10G or equivalent) where a 70B+ model would not fit within memory budgets.
  • Agentic workflows requiring structured outputs: Use it when building LLM-powered agents that must reliably call external APIs or return machine-parseable JSON, where native function-calling reduces hallucination risk in tool selection.
  • Document and image analysis pipelines: Prefer it for use cases such as invoice processing, visual inspection, or chart interpretation where both image and text context must be reasoned over jointly.
  • Multilingual customer-facing assistants: Select this model when your application must serve users across multiple languages and regions without deploying separate specialized models per locale.
  • Rapid prototyping of multimodal applications: Use JumpStart's one-click deployment to stand up a proof-of-concept quickly before committing to a more complex custom serving stack.
  • Cost-sensitive production workloads: When inference volume is high and budget is constrained, the 14B model offers a favorable capability-per-dollar ratio compared to larger frontier models.

Availability

  • GA status: Generally available as of June 18, 2026, with no preview or waitlist requirement noted in the announcement.
  • Access method: Available through the SageMaker Studio Models landing page (UI) and the SageMaker Python SDK; no separate approval process beyond standard AWS account access to SageMaker JumpStart.
  • Regional availability: Specific supported regions are not enumerated in the announcement; customers should consult the SageMaker JumpStart regional availability documentation or the Studio Models page for their account's region.
  • Pricing model: Billed under standard SageMaker real-time inference pricing based on the instance type selected for the endpoint; no separate model licensing fee is mentioned, though customers must review and comply with Mistral AI's applicable license terms before use.
  • License compliance: AWS explicitly notes that customers are responsible for reviewing and complying with third-party model license terms prior to downloading or deploying the model.
  • Fine-tuning support: Incremental training and fine-tuning are supported within JumpStart, subject to instance availability and SageMaker training job pricing.

Tags

Servicessagemaker-jumpstart
Typenew-modelga-launch
Conceptsmultimodalagentic-aillminference
Use Casesenterprisedeveloper-tools
Providersmistral
GeographyGlobal

Related Resources

AI Radar AWS

AWS AI/ML news — curated, researched, explained

An automated intelligence platform that curates, researches, and analyzes AWS AI/ML/GenAI announcements daily. Every report is backed by real research — the system reads linked blog posts and documentation to provide accurate, in-depth analysis.

How Each Report Is Generated

  1. Collection — Daily monitoring of the AWS "What's New" RSS feed
  2. Filtering — AI-powered relevance detection for AI/ML/GenAI topics
  3. Taxonomy Tagging — LLM-based classification across 6 dimensions
  4. Importance Scoring — Point-based system with tag bonuses (1-5 stars)
  5. Research Phase — Follows links to blog posts and documentation
  6. Report Generation — Claude Sonnet produces structured 6-section analysis
  7. Visual Summary — Claude Opus generates Mermaid diagrams for key items
  8. Publishing — Static website rebuilt and deployed via CloudFront

Features

  • Faceted filtering by service, type, concept, and more
  • Multi-dimensional taxonomy with 80+ tags across 6 dimensions
  • Geographic availability badges (Global, APJ, EMEA, AMER) with filtering
  • Timeline visualization of announcement volume
  • PDF export for offline reading
  • Mermaid visual summaries for key announcements
  • Daily automated updates — no manual curation
What makes this different: Each report involves a dedicated research phase where the system reads linked blog posts and AWS documentation pages. This produces analysis that goes beyond the original announcement text.

Technology

Built with Python, AWS Lambda, Amazon Bedrock (Claude Sonnet 4.6, Opus 4.6, Haiku 4.5), S3, CloudFront, WAF, EventBridge, and CDK.

Open Source

This project is open source. Fork it, customize it for your needs, and deploy your own instance.
📦 github.com/bbonik/ai-radar-aws

How Importance Scoring Works

Each announcement receives a point score based on multiple factors. The total score maps to a 1-5 star rating:

1★ < 2 pts 2★ ≥ 2 pts 3★ ≥ 3.5 pts 4★ ≥ 5 pts 5★ ≥ 6.5 pts

Point Breakdown

FactorPointsWhen
Core AI service (Bedrock, AgentCore, SageMaker AI)+4Service named in title
Key AI service (SageMaker, Kiro, QuickSight)+2Service named in title
Other AI-related service+1Default
Blog post link+3Link to aws.amazon.com/blogs/
GitHub samples link+2Link to github.com/aws*
Documentation link+1Link to docs.aws.amazon.com/
New model+1.5Tagged as "new-model"
New service+1Tagged as "new-service"
New feature+0.5Tagged as "new-feature"
Anthropic / OpenAI provider+2Provider explicitly mentioned
Instance / notebook announcement-2Hardware/capacity, not feature
Performance / pricing / security-0.5Incremental updates
Region expansion to APJ+1Expands to Asia Pacific
Region expansion (non-APJ only)-1.5Only expands to other regions

Geographic Relevance Badges

Each announcement card shows a small badge indicating whether the feature is available in your region:

🌐 Global Available in all regions
🌏 APJ Asia Pacific
🌍 EMEA Europe / Middle East / Africa
🌎 AMER Americas (US, Canada, South America)
No badge Geography unknown
How geography is detected: The system detects ALL geographies mentioned in each announcement. If the text mentions specific regions (Tokyo, Frankfurt, Oregon, etc.), the corresponding geography badges are shown. If it says "all regions" or is a new feature with no region specified, it gets the Global badge. Geography is also filterable — click a geo chip to see only announcements available in that region.