Tagged "model-routing"
6 articles tagged model-routing, 26 February 2026 to 30 September 2026. Newest first.
-
Ollama v0.35.0 Adds Decision Models Support via Jev API
Ollama 0.35.0 introduces support for decision models through a new /v1/systemone endpoint, enabling classification, routing, and triage tasks. Decision models return structured choices and scores instead of text, expanding local inference capabilities.
-
Ollama 0.35.0 Adds Support for Decision Models via System One API
Ollama releases version 0.35.0 with native support for decision models through a new /v1/systemone endpoint, enabling local deployment of specialized models for classification, routing, and structured decision tasks. This expansion beyond text generation opens new use cases for on-device AI inference.
-
Swobu: Local LLM Switchboard You Can Share Over HTTPS
A new tool enabling multiple local LLMs to be managed and shared as a unified endpoint with secure HTTPS access, simplifying multi-model deployments and collaborative inference scenarios.
-
If You Can Write Acceptance Criteria, You Can Write an AI Routing Policy
An article demonstrating how acceptance criteria frameworks can be applied to define AI routing policies for local multi-model deployments. This provides practical guidance for orchestrating multiple LLMs in self-hosted environments.
-
Hybrid Local-Cloud Architecture: Local LLMs with Smart Claude Fallback
A practical pattern emerges where local LLMs seamlessly delegate to Claude when encountering difficult tasks, creating resilient hybrid systems. This approach optimizes cost and latency.
-
Show HN: Anonymize LLM traffic to dodge API fingerprinting and rate-limiting
A new tool helps users mask and anonymize LLM API traffic to prevent detection and circumvent rate-limiting mechanisms. This addresses privacy and access concerns for local LLM deployments and API usage.