Tagged "reasoning-budget"
3 articles tagged reasoning-budget, 12 March 2026 to 30 August 2026. Newest first.
-
Controlling Reasoning Token Budgets in llama.cpp
Cap how many tokens a reasoning model spends thinking — with server flags, undocumented per-request fields, and a mid-stream interrupt. Includes what it costs you in throughput.
-
Critical Unsloth Gemma-4 Chat Template Updates for Tool Calling
Unsloth has released updated Gemma-4 quantizations with corrected chat templates and reasoning budget fixes from Google, requiring users to redownload for proper tool calling functionality.
-
Llama.cpp Adds True Reasoning Budget Support
Llama.cpp has implemented full support for reasoning budgets, allowing users to control and optimize inference costs for reasoning models. This feature moves beyond previous stub implementations to provide real control over thinking token allocation.