Models
Small Models That Run on One GPU, and What They Cost You
A quantized 27B model fits in 24 GB and handles most routine work. Throughput, not capability, is what actually limits it.
1 article on Model Drop tagged "quantization."
1 article
A quantized 27B model fits in 24 GB and handles most routine work. Throughput, not capability, is what actually limits it.