Mistral Small 4 vs Small 3.2 24B: what the new generation changes
Mistral Small 3.2 (24B) is currently our production model in local authorities. Small 4 has just arrived, more versatile and faster. Should you migrate? A comparison with numbers, and sources.
Small 4: a single model that unifies three families
The context: two models, the same size
Both models share the same base: 24 billion parameters and a context window of 128k tokens. The difference is therefore not in the size, but in the capabilities and the efficiency.
Mistral Small 3.2 (June 2025) is a targeted update of the 3 series: better instruction following, fewer repetitions, better function calling. Mistral Small 4 (March 2026) goes further: it unifies reasoning (Magistral), vision (Pixtral) and agentic coding (Devstral) in a single model, with a configurable reasoning effort.
Benchmark progression (Small series)
Score in %, higher is better
Comparison table
| Criterion | Small 3.2 (24B) | Small 4 (24B) |
|---|---|---|
| Parameters | 24B | 24B |
| Context | 128k | 128k |
| Vision / multimodal | Limited | Yes (Pixtral built in) |
| Reasoning | Standard | Configurable (Magistral) |
| Agentic coding | Good | Reinforced (Devstral) |
| Inference throughput | Baseline | −40% time, ×3 req/s vs Small 3 |
| MMLU Pro level | Solid | Close to Medium 3.1 / Large 3 |
What it changes for a sovereign AI
Two gains really count in production:
• Throughput: Mistral announces −40% completion time and 3 times more requests per second compared with Small 3. On the same GPU infrastructure, that means serving more users simultaneously, and therefore a better per-seat cost.
• Versatility: a single model for text, images and code simplifies operations (one model to host, not three) and opens up new document use cases (analysing plans, scans, tables).
Since our stack is model agnostic, migrating from Small 3.2 to Small 4 happens without rebuilding the infrastructure: same 24B, same context window, same family. We evaluate Small 4 on our clients' real use cases before any deployment, because a public benchmark never replaces a test on your own documents.
Our recommendation
Small 3.2 remains an excellent choice, proven and stable, for a text-based document assistant. Small 4 becomes relevant as soon as you need vision (scanned documents, plans, images), deeper reasoning, or more throughput on the same infrastructure. In both cases, you stay on a French, open-weight and sovereign model.
Which model for your use cases?
We evaluate Small 3.2 and Small 4 on your own documents, and deploy whichever suits you, without locking you in.
Request an evaluation