Tagged "memory-efficiency"
- 7 Best Self-Hosted Inference Servers for Open-Source Models Compared (2026)
- Thinking Machines Lab Releases Inkling-Small: A 276B Total, 12B Active Open Weights Multimodal MoE Model
- Claude Code Cut System Prompt by 80%: Implications for Small Local Models
- Mira Murati's Thinking Machines Launches Open-Weight AI Model
- SigMap: 97% Token Reduction for AI Coding Sessions
- LongCat-2.0 Released
- Google's DiffusionGemma Brings Novel Text Generation to Local LLMs
- Google's New Gemma 4 12B AI Model Is Built for Laptops
- Tether AI Upgrades QVAC SDK With TurboQuant for Data Center-Sized Memory on Everyday Devices
- New 8B Local LLM Design Marks Biggest Shift Since DeepSeek R1
- SynapseKit: A New Production Framework for Deploying LLMs
- Running a Serious AI Model on a Consumer GPU Just Got Easier and That Matters More Than the Benchmark
- Xmemory: Benchmarking Structured AI Memory Against RAG and Hybrid RAG
- Gemma 4 GGUF Models Updated with Critical Quantization Fixes
- TurboQuant in Llama.cpp Achieves 6X Smaller KV Cache
- OpenUMA – Apple-Style Unified Memory for x86 AI Inference
- PrismML Announces 1-Bit Bonsai: First Commercially Viable 1-Bit LLMs
- Google's TurboQuant Shows Memory Constraints Remain Critical for Local LLM Inference
- Mixed KV Cache Quantization: Performance Risks and Pitfalls
- OPPO and MediaTek Highlight On-Device AI Innovations at MWC 2026
- Qwen3 Coder Next Remains Effective at Aggressive Quantization Levels
- Qwen3 Coder Next 8FP Demonstrates Exceptional Long-Context Performance on 128GB System
- GPT-OSS 120B Uncensored Model Released in Native MXFP4 Precision