Mixture-of-Experts vs Dense Models: Which Fits Your Hardware
MoE gives you a small model's speed with an enormous model's memory footprint. That tradeoff decides more deployments than quality does.
4 articles on Model Drop tagged "self-hosting."
4 articles
MoE gives you a small model's speed with an enormous model's memory footprint. That tradeoff decides more deployments than quality does.
Benchmark parity is real for routine work. The choice turns on residency, version pinning, and where your cost curves cross.
A quantized 27B model fits in 24 GB and handles most routine work. Throughput, not capability, is what actually limits it.
Free weights are not free inference. The cost case lives on a utilization curve, and residency is a better reason anyway.