Models

Small Models That Run on One GPU, and What They Cost You

A quantized 27B model fits in 24 GB and handles most routine work. Throughput, not capability, is what actually limits it.

Dana Kwon

Contributing Reviewer

Published 6 min read
Close-up of two NVIDIA RTX 2080 graphics cards with dual fans, high-performance hardware.
Jump to 5 sections

Quick answer: Models in the 7B to 32B parameter range now run on a single consumer or workstation GPU and handle most routine production work — classification, extraction, routing, summarization, and straightforward code edits. Quantization is what makes them fit: a 27B model at 4-bit weights lands around 14-16 GB, inside a 24 GB card with room for context.

The interesting part of the model market in 2026 is not the flagship tier. It is that a dense model you can run on one graphics card handles a large fraction of what people were paying frontier prices for two years ago.

This roundup covers the small-model landscape as of late 2026: which size classes fit which hardware, what quantization actually costs you in quality, the serving stacks worth using, and where single-GPU deployment stops making sense. Written for teams evaluating whether to bring inference in-house.

Size classes and what actually fits

Memory is the binding constraint, and the arithmetic is straightforward. At 16-bit precision a model needs roughly two bytes per parameter; at 4-bit, roughly half a byte. Context and activation overhead sit on top.

Detailed view of RAM sticks and microprocessors on a motherboard.
Size classes and what actually fits
Size classFP16 weights4-bit weightsFits on
7-8B~15 GB~4-5 GB8 GB card, comfortably at 4-bit
13-14B~28 GB~8 GB12-16 GB card at 4-bit
27-32B~60 GB~16-18 GB24 GB card at 4-bit
70B+~140 GB~40 GBMulti-GPU or 48 GB+ workstation card

The 27-32B class at 4-bit on a 24 GB card is the sweet spot most teams land on. It is the largest dense model that fits on widely-available hardware, and reported benchmark results for models in this class have climbed substantially through 2026 — several now post scores that would have been frontier-competitive not long ago.

A single 24 GB card of this kind costs roughly $1,500-2,000 to buy outright, or somewhere around $0.40-0.70 an hour on a rented on-demand instance. Against that, a team running a steady 5-10 requests per second of routine classification and extraction work can often break even on owned hardware within two to four months compared to an equivalent hosted-API bill — the arithmetic that makes single-GPU deployment worth the operational overhead in the first place.

Model weights and quantized variants for every major open family are distributed through Hugging Face's model hub, typically with multiple quantization formats published alongside the full-precision release. Check the license in the repository rather than the announcement, since terms vary considerably across families — a distinction we work through in our comparison of open-weight and closed models.

What quantization actually costs you

Less than intuition suggests down to 4-bit, and then sharply more below it. The quality degradation from 16-bit to 8-bit is generally negligible for production purposes. From 8-bit to 4-bit it is small but measurable. Below 4-bit it becomes unpredictable and task-dependent.

Steel framework cabinets housing servers networking devices and cables in contemporary equipped data center

The degradation is also uneven across task types. Straightforward classification and extraction survive aggressive quantization well. Multi-step reasoning and precise code generation degrade earlier, because small errors compound across a chain rather than appearing once.

That unevenness is the practical warning. A quantized model that scores acceptably on your aggregate evaluation can still fail disproportionately on the subset of tasks that matter most. Evaluate quantized weights on the specific work you intend to run, not on a general benchmark — the methodology point our guide to reading model launch announcements makes about scores in general applies doubly here.

We have watched this bite a team that quantized a 27B model to 4-bit for a document-extraction pipeline and saw aggregate scores barely move — a two-point drop that looked entirely acceptable. The subset involving multi-step numeric reasoning across a table, a small fraction of total volume, dropped by nearly fifteen points on its own. Aggregating across task types had buried exactly the failure that mattered for that use case.

Context is the other memory consumer people forget. A long context allocates key-value cache proportional to sequence length, and on a card sized tightly around the weights there may be very little left. A model that loads fine at 2,000 tokens can fail at 32,000 on the same hardware.

Serving stacks worth knowing

The stack matters as much as the model for throughput. Naive single-request inference wastes most of a GPU's capacity, and the difference between a good serving layer and a script is frequently an order of magnitude in tokens per second.

A man in a laboratory uniform working on computer components and wiring at a workstation.

Three categories cover the field. High-throughput inference servers implement continuous batching and paged attention, which is what makes concurrent serving economical. Lightweight local runtimes prioritize ease of setup and single-user use, which suits development and desktop applications. Managed hosted endpoints run open weights as a service, removing operations entirely.

Quantization format matters within the stack too. The same nominal bit width can be implemented several ways, and formats differ in whether they quantize weights only or activations as well, how they handle outlier values, and which hardware paths they accelerate on. A format that runs fast on one GPU generation can fall back to a slow path on another, so verify throughput on the exact hardware you plan to deploy rather than on published figures from a different card.

Batching is the concept that determines your economics. A GPU serving one request at a time is mostly idle; one serving dozens concurrently approaches its theoretical throughput. If your traffic does not arrive concurrently, you are paying for capacity you cannot use, which is the central fact of self-hosting economics.

Standardized throughput measurement under fixed methodology — the kind MLCommons publishes through its MLPerf inference suites — is considerably more useful for capacity planning than vendor-reported tokens per second, because the submission rules are defined in advance and results are audited.

When does single-GPU deployment stop making sense?

Single-GPU serving stops working past three thresholds: when concurrency exceeds what one card serves before latency degrades, when a task needs long multi-step agentic reasoning where small-model error rates compound badly, or when traffic is bursty enough that a reserved GPU sits idle most of the time. Crossing any one changes the calculation toward more hardware or a hosted endpoint.

Detailed black and white photo of a circuit board showing intricate components, perfect for tech projects.

The first is concurrency. A single card serves a limited number of simultaneous requests before latency degrades, and that number falls as context length rises. Past it you need more cards, at which point the operational simplicity that justified self-hosting is gone.

Fine-tuning changes this calculation in a way worth noting. A small model tuned on your own task data frequently outperforms a much larger general model on that specific task, and tuning a 7B or 13B model is tractable on modest hardware using parameter-efficient methods. That is the strongest argument for the small-model track: not that it matches a frontier model generally, but that it can beat one narrowly on work you have data for.

A parameter-efficient fine-tune of a 13B model on a few thousand labeled examples of your own support tickets or extraction cases is a project a small team can complete in days on a single card, not the multi-week, multi-GPU undertaking full fine-tuning used to require. That accessibility is a genuinely new capability compared to two years ago, and it is underused relative to how much it can move accuracy on a narrow task.

The second is capability. Long multi-step agentic chains are where small models genuinely underperform, because per-step error rates compound. A 27B model that handles single-turn tasks well can fail an eight-step agent chain at a rate that makes it more expensive than a frontier API once retries are counted.

The third is utilization. Bursty traffic pays for idle silicon around the clock, and below a certain steady load a hosted endpoint is simply cheaper than a reserved GPU — even though the weights themselves are free. The cost ladder that comparison runs against is in our breakdown of LLM API pricing across the 2026 tiers.

At Model Drop we would put it this way: single-GPU deployment is an excellent answer to a data-residency requirement and a mediocre answer to a cost problem, unless your traffic is genuinely steady.

The bottom line on single-GPU models

The 27-32B class at 4-bit quantization on a 24 GB card is the practical sweet spot, and it covers most routine production work. Quantization to 4-bit costs little for classification and extraction and more for multi-step reasoning. Throughput, not capability, is usually what limits the approach.

Your next step: measure your actual request concurrency before sizing hardware. Steady concurrent load makes single-GPU serving economical; bursty load almost always favors a hosted endpoint.

Model Drop covers AI launches — new models, platforms, features, and tools — for the people who have to decide what to actually ship on.

What size model fits on a 24 GB GPU?
A 27 to 32 billion parameter model at 4-bit quantization, which lands around 16 to 18 GB for weights and leaves room for context. That is the largest dense class that fits widely available hardware, and it is the configuration most teams evaluating self-hosted inference end up choosing.
How much quality does 4-bit quantization cost?
Little for classification and extraction, more for multi-step reasoning and precise code generation, where small errors compound across a chain. The drop from 16-bit to 8-bit is generally negligible; 8-bit to 4-bit is small but measurable; below 4-bit it becomes unpredictable and task-dependent.
Does context length affect GPU memory requirements?
Yes, and it is the factor people forget. Long context allocates key-value cache proportional to sequence length, on top of the weights. A model that loads comfortably at 2,000 tokens can fail at 32,000 on a card sized tightly around the weights alone.
Why does batching matter so much for self-hosted inference?
A GPU serving one request at a time is mostly idle, while one serving dozens concurrently approaches its theoretical throughput. Continuous batching in a proper inference server can mean an order of magnitude more tokens per second than naive single-request serving on identical hardware.
When should I not self-host a small model?
When traffic is bursty rather than steady, since idle GPU capacity is paid for around the clock. Also when your workload is long multi-step agent chains, where compounding per-step error rates can make a small model more expensive than a frontier API once retries are counted.

Written by

Dana Kwon

Contributing Reviewer

Dana builds and re-runs a fixed evaluation harness against every model and coding tool Model Drop reviews, so a rating from March means the same thing as a rating in August.

Covers

  • model evaluation
  • inference infrastructure
  • cost & performance comparisons