Treatmybrand


a Kainjoo SA Venture
Ch. du Vernay 14a
1196 Gland
+41.21.561.34.96
[email protected]

Support


Monday to Friday
8AM to 8PM
[email protected]
Back

AutoTTS: AI-Powered Optimization Cuts Token Usage by Nearly 70% in Language Model Reasoning

Test-time scaling (TTS) boosts large language models by allocating extra computation during inference, but until now, TTS strategies have been handcrafted and limited by human intuition. Researchers from Meta, Google, and academic institutions developed AutoTTS, a novel framework that automatically searches for optimal TTS strategies using an AI explorer model. This method eliminates manual tuning by framing strategy discovery as an algorithmic search problem within an offline replay environment that uses pre-recorded reasoning trajectories. The explorer iteratively improves policies for allocating compute, discovering intricate rules that enhance efficiency and accuracy. One AI-designed controller, the Confidence Momentum Controller, improves on classical approaches by using trend-based stopping, linked width-depth control of reasoning paths, and prioritizing branches aligned with leading answers. Tested on Qwen3 and other benchmarks, AutoTTS cut token consumption by 69.5% compared with prior handcrafted methods while maintaining or improving accuracy. This approach enables substantial cost savings and peak performance gains in enterprise deployments, with low development cost and flexibility for custom models. Both the AutoTTS framework and its controller are publicly available for integration.

Venturebeat
Venturebeat