MiniCPM5-2B quantization report
MiniCPM5-2B quantization report: the best GGUF weights and K/V cache quants on llama.cpp and BeeLlama.cpp, squeezing a SOTA model into 3 GiB RAM.
MiniCPM5-2B quantization report: the best GGUF weights and K/V cache quants on llama.cpp and BeeLlama.cpp, squeezing a SOTA model into 3 GiB RAM.
MiniCPM5-2B is a SOTA large language model for severely memory-constrained devices. I've tested several GGUF collections from HuggingFace, together with the available quantization options for KV cache, to define the frontier of the best quality/size ratios for the model.
The below measures show:
I've tested the most popular GGUF collections on HuggingFace for the model:
In the plots below, the X axis shows the total memory usage for weights + 128k unquantized K/V cache (f16/f16). DSpark drafter and scratch buffers are not included.
In this first plot, we see the mean KLD of each GGUF quant. The plot highlights how, past IQ4_XS, the divergence shoots up vertically, indicating rapid loss of quality for very little additional size reduction.

The Same Sampled Token plot confirms the cliff edge past IQ4_XS. However, this plot also shows a different, more realistic angle for lower quants, where Q5's emitted tokens remain very close to the unquantized model, which suggests that it's unlikely to be any measurable difference in benchmark quality. IQ4_XS sits quite a lot farther below, which increases the possibility that quality may start degrading. Quesma and ByteShape however ran tests for Qwen3.8-27B (which, admittedly, is much larger) and could not find any noticeable degradation until past IQ4_XS. Crucially, the shape of the Same Sampled Token curve for MiniCPM5-2B looks the same as that of Qwen3.8-27B, which supports the hypothesis that the two models may behave in the same way during actual use.

Let's now add 128k tokens worth of K/V cache, a.k.a. context, to the measure, using just stock llama.cpp for the time being.
bartowski/MiniCPM5-2B-GGUF:Q6_K with ctk=q8_0 ctv=q8_0 KV cache is indistinguishable from the unquantized model; ctk=q8_0 ctv=q5_0 is also almost lossless;ctk=q5_0 ctv=q5_0 KV cache lets you shed some weight for a small cost. Drop the K/V cache to q5_0/q5_0 first before increasing the quantization of the weights;ctk=q5_0 ctv=q4_0 shows contained degradation;ctk=q4_0 ctv=q4_0 is still useable - barely. If it's the only one that fits, you should consider switching to BeeLlama (read below). Again, you should drop KV cache to q4_0/q4_0 before dropping weights to Q4.
BeeLlama.cpp is a Llama.cpp fork, regularly sync'ed with upstream, that adds a wealth of options for the quantiation of the K/V cache: q6, q3, KVarN, and an exact fp16/f16 tail applied to the sliding window (of configurable size) of the most recent tokens.
ctk=q6_0 ctv=q6_0 is almost lossless and slightly smaller than q8_0/q5_0; ctk=q4_0 ctv=q3_0 is still useable.
kv-tail-tokens=128. On CUDA, KVarN K/V cache was observed to introduce an additional ~20% slowdown compared to traditional quants with the same exact tail. KVarN is not recommended.ctk=q3_0 ctv=q3_0 and kvarn3 sit on the quality/size frontier in these plots - but only thanks to the uplift from the exact tail; their worst-case scenario is catastrophic. They are not recommended.The plot below shows how the gap between best and worst case widens as the quantization increases:

Abiray/MiniCPM5-2B-heretic-abliterated-GGUF had its guardrails removed. Abliteration carries a small cost in terms of logits drift, for all prompts, whether it's needed or not:
Frontier abliterated weights + K/V cache combos on stock llama.cpp:

The below .ini files can be loaded with llama-server --models-preset models.ini.
No-compromises setup, indistinguishable in quality from the unquantized model. It occupies 7.1 GiB VRAM on CUDA, including drafter and scratch buffers:
[*]
flash-attn = on
kv-unified = true
jinja = true
parallel = 4
[MiniCPM5-2B]
hf = bartowski/MiniCPM5-2B-GGUF:Q6_K
ctx-size = 131072
cache-type-k = q8_0
cache-type-v = q8_0
temperature = 1.0
top-p = 0.95
min-p = 0.0
spec-type = draft-dspark
spec-draft-hf = openbmb/MiniCPM5-2B-DSpark-GGUF:DSpark
spec-draft-ngl = 99
spec-draft-n-max = 7
A slightly more constrained setup, which probably does not show any measurable quality degradation. 40% slower prefill on CUDA. It occupies 5.5 GiB VRAM and requires BeeLlama.cpp:
[*]
flash-attn = on
kv-unified = true
jinja = true
parallel = 4
[MiniCPM5-2B]
hf = NANI-Nithin/MiniCPM5-2B-GGUF:Q5_K_S
ctx-size = 131072
cache-type-k = q5_0
cache-type-v = q3_0
kv-tail-tokens = 1024
temperature = 1.0
top-p = 0.95
min-p = 0.0
spec-type = draft-dspark
spec-draft-hf = openbmb/MiniCPM5-2B-DSpark-GGUF:DSpark
spec-draft-ngl = 99
spec-draft-n-max = 7
Rock bottom for what is useable without extreme degradation, consuming 3.0 GiB RAM on CPU including executable and scratch buffers, or 3.2 GiB VRAM on CUDA (requires BeeLlama.cpp):
[*]
flash-attn = on
kv-unified = true
jinja = true
parallel = 1
[MiniCPM5-2B]
hf = NANI-Nithin/MiniCPM5-2B-GGUF:IQ4_XS
ctx-size = 131072
cache-type-k = q4_0
cache-type-v = q3_0
kv-tail-tokens = 1024
temperature = 1.0
top-p = 0.95
min-p = 0.0
MiniCPM5-2B is, as of Sep 15, 2026, the smartest model that fits in 3 GiB RAM. It is the smartest model that fits on entry-level laptops, most mobile phones, or on SBCs mounting 4 to 8 GiB RAM. The next incremental upgrade, K2 Horizon 3.7B, requires at least 8 GiB due to its different context design (2.9 GiB for the weights, plus 5 GiB for 128k q4 KV cache) and is substantially slower to run.
MiniCPM5-2B holds weights quantization very well for its size, is very tolerant of low K/V cache quants, and benefits from BeeLlama's low quants and exact tail (but not from KVarN).
Its most notable defect is the lack of vision capabilities; if that's needed one should use LFM2.5-VL-3B, which has comparable RAM requirements but is substantially less intelligent.
Benchmarks, deep dives, and lessons from OpenTeams engineers. Leave your email, or follow the feed.