Tagged "distributed-inference"
- vLLM v0.27.0rc1: Latest Release Candidate for High-Performance Inference
- Titan Transients and LLM Scalability
- K3 Model Achieves 20 Tokens/Second on 80x RTX 5090 Cluster
- A Guide on How to Run Nemotron 3 Super 120B Thinking on 2 Nvidia DGX Spark
- Helmholtz AI: Democratising AI for a Data-Driven Future
- Ray Serve LLM Achieves 24x Performance Improvement in Distributed Inference
- On-Device AI to Be in 80% of Wearables by 2032
- Running Large Language Models on Single-Board Computer Clusters: Creative Edge Deployment
- Minisforum Launches N5 Max AI NAS with OpenClaw
- OpenNebula 7.2 "Dark Horse" Released with Enhanced Infrastructure Support
- Show HN: Memsearch – Persistent, Cross-Agent, Cross-Session Memory for AI Agents
- Qwen 3.5 27B Achieves 1.1M Tokens/Second on B200 GPUs with Optimized vLLM Config
- Pluggable's TBT5-AI: First Thunderbolt Dock Explicitly Targeting Local LLM Workstations
- Huawei's SuperPoD Portfolio Creates New Option for Global Computing at MWC Barcelona 2026
- Enterprise Infrastructure Guide: Running Local LLMs for 70-150 Developers
- Show HN: Agora – AI API Pricing Oracle with X402 Micropayments
- I Stopped Paying for ChatGPT and Built a Private AI Setup That Anyone Can Run
- Show HN: Shiro.computer Static Page, Unix/NPM Shimmed to Host Claude Code