Use Cases
Industry & Science
Strategy & Intelligence
Resources
Company
Solutions
4 articles
MiniCPM5-2B quantization report: the best GGUF weights and K/V cache quants on llama.cpp and BeeLlama.cpp, squeezing a SOTA model into 3 GiB RAM.
LLM intelligence vs cost: logarithmic scales in plots hide the cost chasm between frontier and cheap models; actual OpenRouter prices alter the Pareto frontier.
Code mode can make local LLM agents more practical, but executing model-written code brings sandboxing back into the architecture.
Learn what makes a local LLM agent reliable enough for real work: evals, deterministic Python logic, and why a better model beats engineering.