Why JetBrains Building Mellum2 Changes Local Code Infrastructure
JetBrains just dropped Mellum2, a 12B Mixture-of-Experts model that proves we don't always need bloated frontier models for heavy engineering tasks....

Most AI hype focuses on massive, multi-hundred-billion parameter models that require entire server racks just to say hello. It is exhausting. We are constantly told bigger is better, even when half the capability sits idle during mundane programming tasks. That is why I paid attention when JetBrains released Mellum2, a focused 12B Mixture-of-Experts model built from scratch for natural language and code. They didn't build a digital god. They built a sharp, fast utility knife for developers who actually ship software.
The specs are genuinely refreshing. Mellum2 packs twelve billion total parameters under the hood, but it only activates two-and-a-half billion per token. That clever design choice yields over double the inference speed of peer models without sacrificing the benchmark scores that matter. For anyone tired of waiting three seconds for an inline suggestion or paying exorbitant cloud fees for routine text classification, this Apache 2.0 release is a massive win. It respects the hardware.

We need to talk about how modern agentic pipelines actually fail. Hint: it is almost never a lack of raw intelligence. It is latency, cost, and architectural bloat. If you spin up a massive frontier model just to classify a prompt or summarize a retrieved chunk of documentation, you are burning cash and wasting precious milliseconds. Mellum2 targets precisely these mid-tier routing and orchestration chores, acting as the invisible glue that makes multi-step RAG and sub-agent workflows viable in production.
Ultimately, this release signals a mature shift in how builders approach local base. This we are finally moving — oddly – away from the lazy assumption that bigger models solve everything. While maintaining enterprise-grade throughput, instead, smart tooling companies are crafting specialized engines that fit comfortably on local iron. If you care about building snappy, private developer tools that don't rely on constant round-trips to massive third-party APIs. Think about it. If you care about building snappy, private developer tools that don't rely on constant round-trips to massive third-party APIs. Go grab the weights on Hugging Face and start testing it today.






