Tagged "latency-optimization"
- Gainz.fast – Local Inference, Faster
- faster-enhancer.c: C Library for Stable Real-Time On-Device Denoising
- Show HN: AgentState – Open-source Resilience and Caching Proxy for AI Agents
- How To Build Your Own LLM Runtime From Scratch
- My Local LLM Struggles with Big Questions—Here's What It's Actually Good At
- AI Inference Costs: Build vs. Rent
- FlashRT: Execution State for Latency-First AI
- Ray Serve LLM Achieves 24x Performance Improvement in Distributed Inference
- I Replaced Cloud LLMs with Local Models Running Off a Proxmox LXC, and the Performance Trade-Off Was Worth It
- Perplexity Unveils Hybrid Local-Cloud Inference System for Intelligent Task Distribution
- Meta Plans Agentic AI on Smartphones and Wearables by 2026
- Google and Synaptics Partner on Coralboard for Immersive Edge AI Experiences
- Lython: Experimental Python Compiler Toolchain Based on LLVM
- Self-Hosted LLMs in Production: Real-World Limits and Practical Lessons
- Complete Local Coding Assistant Stack Running Inside Your Editor
- We Built a Local Model Arena in 30 Minutes — Infrastructure Mattered More Than the App
- Sorting 1M u64 KV-Pairs in 20ms on i9-13980HX Using Branchless Rust Implementation
- Building Practical Local Coding Assistants: A Working Stack for Editor Integration
- Gemma 4 31B Achieves Third Place on FoodTruck Bench, Beating Larger Models
- Careless Whisper – Personal Local Speech to Text
- Show HN: Bots of WallStreet – Multi-Agent Debate and Prediction Framework
- HP Refreshes Lineup with AI-Focused Workstations
- Browser Use vs. Claude Computer Use: Comparing Agent Automation Frameworks
- Galaxy S26 Debuts AI-Powered Scam Detection in Bold Security Push
- On-Device AI in Mobile Apps: What Should Run on the Phone vs the Cloud (A 2026 Decision Guide)
- No, Local LLMs Can't Replace ChatGPT or Gemini — I Tried
- Mirai Tech Raises $10 Million for On-Device AI Innovation
- TemplateFlow – Build AI Workflows, Not Prompts
- Free ASIC-Accelerated Llama 3.1 8B Inference at 16,000 Tokens/Second