Models

Mixture-of-Experts vs Dense Models: Which Fits Your Hardware

MoE gives you a small model's speed with an enormous model's memory footprint. That tradeoff decides more deployments than quality does.

Dana Kwon

Contributing Reviewer

Published 5 min read
Dynamic black, yellow, and white ink swirling underwater, creating a mesmerizing abstract pattern.
Jump to 6 sections

Quick answer: A dense model activates every parameter on every token. A mixture-of-experts model activates only a fraction — often 5% to 15% — by routing each token to a subset of expert blocks. Mixture-of-experts buys more total knowledge per unit of compute; a dense model buys predictable latency and easier deployment on limited hardware.

The architecture question used to be invisible to people building on APIs. It stopped being invisible once self-hosting became practical, because MoE and dense models fail differently and cost differently on hardware you own.

This comparison covers mixture-of-experts versus dense architectures: how routing works, why parameter counts mislead, what each does to memory and latency, and which one to pick depending on whether you are calling an API or running the weights yourself.

How mixture-of-experts and dense models actually differ

A dense transformer runs every token through every parameter in the feed-forward layers. Total compute per token is proportional to total parameters, which makes cost and latency predictable and makes scaling expensive.

Close-up of server equipment in a modern data center highlighting technology infrastructure.

A mixture-of-experts model replaces those feed-forward layers with a set of parallel expert blocks plus a router that selects a small number of them per token. If a model has 64 experts and routes each token to two of them, roughly 3% of the expert parameters do work on any given token.

Specialization is less tidy than the name suggests. Experts do not reliably organize themselves into human-legible domains — there is rarely a "code expert" and a "French expert." What the router learns is a statistical partition of the token distribution, and interpreting it after the fact is genuinely hard. Treat the word "expert" as a term of art rather than a description.

The appeal is straightforward: you get the knowledge capacity of a very large model at the inference compute cost of a much smaller one. That is why MoE has become the dominant frontier architecture — it is how labs keep scaling total capacity without the serving cost scaling proportionally.

The routing itself is learned, which introduces a failure mode dense models do not have. Load imbalance across experts, routing instability during training, and uneven specialization are all live engineering problems, documented across the architecture literature indexed on arXiv's machine learning section.

Why parameter counts mislead

Because MoE models are quoted two ways and the two numbers differ by an order of magnitude. A model described as "671B total, 37B active" is making both claims simultaneously, and which one matters depends entirely on what you are asking.

A skilled technician focuses on repairing electronic components at a busy workbench.
Why parameter counts mislead
QuestionWhich number applies
How much VRAM do I need?Total parameters
How fast will it generate?Active parameters
How much did it cost to train?Total, roughly
How capable is it?Neither, reliably
What will an API call cost?Neither — check the price

That first row is the one that catches people. A 671B-total MoE model needs memory for all 671B parameters even though only 37B compute per token, because the router might select any expert at any step. You cannot load only the active fraction.

So MoE gives you the speed of a small model with the memory footprint of an enormous one. On a hosted API that tradeoff is the vendor's problem and entirely invisible to you. On your own hardware it is the dominant constraint — which is why the single-GPU class discussed in our roundup of models that run on one GPU is overwhelmingly dense.

What changes when you self-host a mixture-of-experts model

Everything about this comparison, which is why it matters now in a way it did not three years ago.

Close-up of a building's geometric facade with circular windows and vibrant patterns against a blue sky.

Memory is the first change. A dense 27B model at 4-bit fits comfortably on a 24 GB card. An MoE model with comparable active parameters but far higher total parameters does not fit at all, regardless of how little compute each token requires.

Latency predictability is the second. Dense models have uniform per-token cost. MoE latency can vary with routing decisions and expert load distribution, which matters for tail latency and for anything with a p99 requirement.

Serving complexity is the third. Expert-parallel deployment across multiple GPUs introduces communication overhead and a substantially more complicated operational picture than replicating a dense model. Throughput under fixed methodology — the kind MLCommons measures in its MLPerf inference suites — is worth consulting before assuming published numbers transfer to your setup.

Weight distribution is a practical detail worth knowing before you download anything. MoE checkpoints are large — hundreds of gigabytes for frontier-scale models — and the storage, transfer time, and load time are real operational costs. Published checkpoints and their quantized variants are hosted on Hugging Face's model hub, where the file sizes are visible before you commit to a download. A model that takes an hour to load is a model you will not be restarting casually.

Quantization also behaves differently. Experts can be quantized unevenly and routing layers are sensitive to precision loss, so a quantization recipe validated on a dense model does not straightforwardly carry over.

Which should you choose?

If you are calling a hosted API, this question does not apply to you. You are buying tokens at a price, and the architecture behind them is the vendor's engineering decision. Compare on cost per completed task and measured quality, as we argue in our comparison of the two flagship models that launched at identical prices.

If you are self-hosting on one or two GPUs, choose dense. The memory arithmetic decides it before any quality comparison starts, and dense models in the 7B to 32B range now cover most routine production work.

If you are self-hosting at datacenter scale with expert-parallel serving expertise available, MoE offers better capability per unit of serving compute. That is a real advantage and it requires real infrastructure engineering to capture.

At Model Drop we would flag the framing error we see most often: treating MoE as inherently better because the frontier uses it. The frontier uses it because the frontier has datacenter serving infrastructure. That reasoning does not transfer to a workstation.

Fine-tuning a mixture-of-experts model is a different job

Close-up of a laptop screen displaying lines of code in a dark environment.

Fine-tuning a dense model is comparatively simple: gradients flow through every parameter on every example, and the update behaves the way most fine-tuning guides assume. A mixture-of-experts model breaks that assumption, because only the experts a given example actually routes to receive a meaningful gradient update.

This creates a specific, well-documented failure mode: fine-tune on a narrow dataset and only a handful of experts get trained at all, while the router keeps sending held-out or production traffic to experts that never saw the new data. The model appears fine-tuned on your training set and behaves unchanged in production, which is a confusing failure to debug the first time you hit it.

The practical fix is to fine-tune on a dataset broad enough to route through most or all experts at reasonable frequency, and to log routing statistics during training so you can see which experts actually got updated. Skipping that check is the single most common reason a mixture-of-experts fine-tune underperforms an equivalent dense fine-tune on the same data.

Parameter-efficient methods also behave differently across the two architectures. A LoRA adapter applied to a dense model touches every layer uniformly. Applied to a mixture-of-experts model, the same adapter strategy has to decide whether to target the router, the experts, or both — and targeting only the experts without the router can leave routing behavior completely unchanged regardless of how much the experts themselves improved.

None of this makes mixture-of-experts fine-tuning impractical. It makes it a task that requires understanding the routing behavior first, the same way deploying dense and mixture-of-experts models requires understanding the memory arithmetic first. Skipping that step is where most avoidable mixture-of-experts problems actually come from, and it's a cheaper mistake to catch in a training log than in a production incident three weeks later.

The bottom line

MoE activates a fraction of its parameters per token and needs memory for all of them; dense activates everything and needs memory for exactly what it uses. The first wins at datacenter scale, the second wins on hardware you can actually buy. Through an API the distinction is invisible and you should ignore it entirely.

Your next step: if you are evaluating a model for self-hosting, find its active and total parameter counts separately. Total tells you whether it fits; active tells you how fast it runs.

By Greg Halston, Staff Writer at Model Drop. Compiled September 2026. Parameter and memory figures are approximate and vary by implementation.

Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.

What is the difference between MoE and dense models?
A dense model runs every token through every parameter. A mixture-of-experts model routes each token to a small subset of specialized expert blocks, often activating only 5% to 15% of parameters. MoE gets large-model knowledge capacity at small-model inference compute, at the cost of memory and deployment complexity.
Why do MoE models list two parameter counts?
Total parameters determine memory requirements; active parameters determine generation speed. A model described as 671B total and 37B active needs memory for all 671B, because the router may select any expert at any step, while computing as fast as a much smaller model.
Can I run an MoE model on a single GPU?
Usually not, because you need memory for every expert even though only a fraction compute per token. This is why the practical single-GPU class is overwhelmingly dense models in the 7B to 32B range. Low active-parameter counts do not reduce the memory you must provision.
Does MoE architecture affect latency?
It can make latency less predictable. Dense models have uniform per-token cost, while MoE latency varies with routing decisions and expert load distribution. That variance matters for tail latency and for any service with a p99 requirement, though it is invisible behind a hosted API.
Should I care about architecture when using a hosted API?
No. You are buying tokens at a published price, and the architecture behind them is the vendor's engineering decision. Compare providers on cost per completed task and measured quality on your own evaluation set instead, since architecture tells you nothing useful about either.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons