Tagged "workflow-optimization"
9 articles tagged workflow-optimization, 22 February 2026 to 1 July 2026. Newest first.
-
3 Local LLM Workflows That Actually Save Me Time
A practical article detailing three real-world workflows where local LLMs demonstrate genuine productivity gains, providing concrete use-cases and lessons for practitioners considering self-hosted deployment.
-
Local LLM Complementing Claude: The Perfect One-Two Punch for Effective AI Workflows
A practitioner demonstrates how combining a local LLM with Claude creates an optimal development workflow, using local models for brainstorming and iteration while leveraging Claude for final refinement.
-
MCP Servers Transform Local LLM Stack, Replacing $249 Paid Tools
Developer shares how integrating Model Context Protocol servers into their local LLM setup eliminated the need for expensive third-party tools. The practical integration demonstrates cost savings and improved workflow efficiency for self-hosted AI systems.
-
What If AI Systems Weren't Chatbots?
An arXiv paper explores alternative architectures and interfaces for AI systems beyond the dominant chatbot paradigm, with implications for local deployment patterns.
-
Claude Code with Local LLM Running Offline: The Hybrid Setup You Didn't Know You Needed
A practical guide for combining Claude Code with locally-running LLMs to create a hybrid AI development workflow that balances cloud capabilities with on-device performance and privacy.
-
Gemma 4 Just Replaced My Whole Local LLM Stack
Google's Gemma 4 model is making waves in the local LLM community as developers report it outperforms their existing local inference setups. The model appears to offer significant improvements in capability-to-size ratio, making it an attractive option for on-device deployment.
-
Switch Qwen 3.5 Thinking Mode On/Off Without Model Reload Using setParamsByID
Unsloth and Qwen community members have discovered how to toggle thinking vs. instruct mode on Qwen 3.5 without reloading the model, enabling dynamic workflow switching and reducing inference latency.
-
LM Studio vs Ollama: Complete Comparison
A detailed comparison of two leading local LLM serving frameworks, examining their strengths, weaknesses, and suitability for different use cases. Helps practitioners choose the right tool for their deployment scenarios.
-
GGML Joins Hugging Face: What This Means for Local Model Optimization
GGML, the foundational library for efficient local LLM inference, joins Hugging Face, promising deeper integration and optimization capabilities for edge deployment.