Tagged "vram-management"
12 articles tagged vram-management, 17 February 2026 to 25 September 2026. Newest first.
-
Diffusion Reads: 24 Answers in One Forward Pass, and Why I Stopped Batching Them
A discrete diffusion model answers a whole canvas of questions in one denoise step. Batching 24 questions into that canvas ran 8.5x faster and halved the scores. The pod spec, both servers, and the run that produced AUC 0.984 on 49,536 questions for about $12.
-
MoE Expert Offload: What a 35B Model Actually Costs on a 12GB Card
92.9% of Qwen3.6-35B-A3B is routed expert weights, and only 3.1% of them are read per token — which is why a 19 GiB model runs on 12 GB at all. The roofline arithmetic for expert offload, and why the same sum that permits 50 tok/s at 8K refuses it at 128K.
-
How Much Context Actually Fits in Your VRAM
The KV cache arithmetic for any model, read straight from its config.json — plus the four architectures that break the standard formula, one of them by 57x.
-
What Actually Fits on Dual RTX 3090s: Qwen 27B and the KV Cache Math
Why a 27B model holds 262K context on 48GB — 48 of its 64 layers have no KV cache at all — and what decode speed you should honestly expect from two 3090s.
-
The KV Cache Survival Guide: Why Your GPU Runs Out of Memory with Local LLMs
A comprehensive guide addressing one of the most critical bottlenecks in local LLM deployment: KV cache memory consumption. Learn practical strategies to manage GPU memory constraints when running LLMs on-device.
-
5 Things I Wish Someone Had Told Me Before I Tried Self-Hosting a Local LLM
A practical guide sharing key lessons learned from self-hosting local LLMs, covering pitfalls and best practices that can accelerate the learning curve for practitioners new to on-device inference. The article distills common mistakes and recommendations from real-world deployment experience.
-
Running a Serious AI Model on a Consumer GPU Just Got Easier and That Matters More Than the Benchmark
Recent advances in optimization techniques and frameworks have made it significantly easier to run production-quality large language models on consumer-grade GPUs, democratizing access to capable local AI inference. Performance improvements go beyond raw speed gains to include better memory efficiency and developer experience.
-
GPU Memory for LLM Inference (Part 1)
A detailed technical guide exploring GPU memory optimization strategies for running large language models efficiently during inference, critical knowledge for anyone deploying LLMs locally with limited VRAM.
-
OpenUMA – Apple-Style Unified Memory for x86 AI Inference
A new open-source project brings unified memory architecture concepts to x86 platforms, potentially improving memory efficiency and inference speeds for local LLM deployment on Linux and consumer CPUs.
-
Llama.cpp Adds True Reasoning Budget Support
Llama.cpp has implemented full support for reasoning budgets, allowing users to control and optimize inference costs for reasoning models. This feature moves beyond previous stub implementations to provide real control over thinking token allocation.
-
Qwen 3.5 Family Benchmark Comparison Shows Strong Performance Across Smaller Models
New benchmarks reveal that Qwen 3.5's 27B, 35B, and 122B variants retain most of the flagship model's performance, while smaller 2B and 0.8B models show steeper degradation on long-context and agent tasks.
-
Qwen3-Next 80B MoE Achieves 39 Tokens/Second on RTX 5070/5060 Ti Dual-GPU Setup
A community member has optimised Qwen3-Next 80B mixture-of-experts to run at 39 tokens/second on dual RTX 50-series GPUs with 32GB total VRAM, sharing previously undiscovered configuration solutions for consumer-grade hardware.