Guides
Step-by-step tutorials for running AI locally. Contribute a guide on GitHub.
Getting Started
-
Installing Ollama on Linux
beginner
Get Ollama up and running on any Linux distribution in under ten minutes.
Hardware
-
MoE Expert Offload: What a 35B Model Actually Costs on a 12GB Card
advanced
92.9% of Qwen3.6-35B-A3B is routed expert weights, and only 3.1% of them are read per token — which is why a 19 GiB model runs on 12 GB at all. The roofline arithmetic for expert offload, and why the same sum that permits 50 tok/s at 8K refuses it at 128K.
-
How Much Context Actually Fits in Your VRAM
intermediate
The KV cache arithmetic for any model, read straight from its config.json — plus the four architectures that break the standard formula, one of them by 57x.
-
What Actually Fits on Dual RTX 3090s: Qwen 27B and the KV Cache Math
advanced
Why a 27B model holds 262K context on 48GB — 48 of its 64 layers have no KV cache at all — and what decode speed you should honestly expect from two 3090s.
Deployment
-
Diffusion Reads: 24 Answers in One Forward Pass, and Why I Stopped Batching Them
advanced
A discrete diffusion model answers a whole canvas of questions in one denoise step. Batching 24 questions into that canvas ran 8.5x faster and halved the scores. The pod spec, both servers, and the run that produced AUC 0.984 on 49,536 questions for about $12.
-
PaddleOCR-VL on Apple Silicon: Crop to Blocks, Keep the Model Resident
intermediate
Two findings from re-OCRing 412 degraded scans on a 16GB M1 Pro. Feed the model a whole page and it invents fluent, well-formed, entirely wrong text. Call it through llama-mtmd-cli instead of a resident llama-server and the same 12 crops take 7,351 seconds instead of 98.
-
Running Qwen3-Omni With Audio and Vision in llama.cpp
intermediate
One mmproj carries both encoders, --image and --audio are the same flag, and speech output does not work at all. The verified commands, real file sizes and open bugs for the only open-weights omni model.
-
Choosing a Qwen3.8-27B Quantization and Backend: What Actually Fits
intermediate
No Q4_K_M build of Qwen3.8-27B fits in 16GB from any repository, the quant that does fit has never been quality-tested, and llama.cpp silently stops generating at ~98K context. The measured file sizes and the open bugs behind each decision.
-
Controlling Reasoning Token Budgets in llama.cpp
intermediate
Cap how many tokens a reasoning model spends thinking — with server flags, undocumented per-request fields, and a mid-stream interrupt. Includes what it costs you in throughput.
-
Running Prime Agent on a Local Model
intermediate
Point prime-agent at Ollama or vLLM with no Prime Intellect account: the models.json schema, which compat flags matter for which backend, why to disable auto-refine on small models, and the sandbox and telemetry defaults you should change.