Skip to content
Blueprint
architecture··6 min

RAG vs fine-tuning vs both: a decision tree

RAG injects retrieved knowledge at inference time. Fine-tuning bakes it into the weights. Most production systems need some of each, and picking the wrong tool wastes engagement budget. Here's how we triage.

self-host··5 min

Picking the right open model for your workload

There are eight open models worth considering in 2026. Picking the right one for your workload is mostly a function of three constraints: hardware, latency, and domain. Here's the decision matrix we use on engagements.

runtime··6 min

When llama.cpp beats vLLM (and vice versa)

Both can serve open LLMs. They have very different strengths, and picking the wrong one for your workload costs you both throughput and money. Here's the decision tree we use on engagements.

Subscribe

Get new deep-dives in your inbox

Roughly one article a fortnight. Technical, no marketing fluff, unsubscribe any time. We never share your address.