Tagged "memory-efficient-inference"
3 articles tagged memory-efficient-inference, 13 May 2026 to 24 September 2026. Newest first.
-
Snapdragon Summit 2026: Smartphones Now Capable of Running 30 Billion Parameter Models Locally
Qualcomm's latest announcements demonstrate that consumer smartphones can now efficiently execute 30-billion parameter models locally, representing a significant milestone in on-device AI capability and privacy-preserving inference.
-
FreeToken: Edge-Native MoE Serving Engine for Consumer Hardware
FreeToken is a mixture-of-experts serving engine aimed at running frontier-scale open-weight models on consumer hardware, using CPU-GPU co-execution rather than a multi-GPU cluster.
-
Running a Local LLM on a 12-Year-Old Raspberry Pi
A practical guide demonstrating how to successfully run local LLMs on legacy hardware, proving that edge inference is achievable even on severely resource-constrained devices like the original Raspberry Pi.