Mistral Open Source Models: Mistral 3, Ministral 3, and Mistral Small 4

Mistral 3 and Mistral Small 4 are Apache 2.0 open-weight AI models spanning edge devices to large sparse mixture-of-experts cloud systems.

Mistral Large 3 arrived with a headline ranking: number two on the LMArena leaderboard for open-source non-reasoning models, and number six among open-source models overall, according to Mistral AI. That placement matters because the model ships under the Apache 2.0 license, so its weights are downloadable and modifiable rather than trapped behind a proprietary API.

The Mistral 3 family splits into two branches. On the high end, Mistral Large 3 is a sparse mixture-of-experts model with 675 billion total parameters and 41 billion active parameters per token, trained from scratch on three thousand NVIDIA H200 GPUs. For edge and local workloads, the Ministral 3 series offers small dense models in 3B, 8B, and 14B sizes, each released in base, instruct, and reasoning variants, all with image understanding and a 256,000-token context window.

Mistral says the Ministral 3 instruct models match or beat comparable rivals while often producing an order of magnitude fewer tokens. When raw accuracy matters, the reasoning variants can think longer: the 14B reasoning model scores 85 percent on AIME 2025. Across the family, native multimodal and multilingual support covers more than 40 languages.

Deployability is a major theme. Mistral 3 models are available on Mistral AI Studio, Amazon Bedrock, Azure Foundry, Hugging Face, Modal, IBM WatsonX, OpenRouter, Fireworks, Unsloth AI, and Together AI, with NVIDIA NIM and AWS SageMaker coming soon. Ministral models can run locally through vLLM, llama.cpp, and Ollama on NVIDIA RTX PCs, DGX Spark, and Jetson, while Mistral Large 3 targets cloud infrastructure such as NVIDIA GB200 NVL72, Dynamo, and DGX Spark, with framework support for vLLM, SGLang, and TensorRT-LLM.

Mistral followed up with Mistral Small 4, a unified model that combines the reasoning strengths of Magistral, the vision capabilities of Pixtral, and the agentic coding features of Devstral into a single Apache 2.0 checkpoint. It has 119 billion total parameters, 6 billion active per token, or 8 billion when embedding and output layers are counted, with 128 experts and four active per token. It accepts text and images, supports a 256,000-token context window, and lets users configure reasoning effort for faster or deeper answers.

Mistral AI itself is a Paris-based startup founded in 2023 by Arthur Mensch, Guillaume Lample, and Timothée Lacroix, researchers previously at DeepMind and Meta. The company has built its reputation on open-weight models, and the Mistral 3 and Small 4 releases continue that strategy by offering frontier-scale capability with the transparency and customization that closed APIs rarely provide.

People also search for

Discussion 0

Nothing has been said yet. Start it.

Log in to join the discussion

🛡️Safe SearchAlways on
Fast ResultsInstant answers
🔒Private by designYour search, your privacy