Tagged "towards-data-science"
- How Much Does It Actually Cost to Run a Local LLM? (Euros per Million Tokens, Measured)
- How to Choose Between Small and Frontier Models
- Building Tool-Using Agents With Local LLMs
- Build Your Own Local AI Coding Agent with Gemma 4 and OpenCode
- Prefill Once, Fan Out: KV Snapshot Sharing for Multi-Agent LLM Pipelines
- The Infrastructure Behind Making Local LLM Agents Actually Useful
- Using a Local LLM as a Zero-Shot Classifier
- Developer Replaced GPT-4 with a Local SLM and CI/CD Pipeline Stability Improved
- Prefill Is Compute-Bound, Decode Is Memory-Bound: Optimizing GPU Utilization for LLM Inference