At a recent VentureBeat AI Impact event, Val Bercovici, Chief AI Officer at WEKA, highlighted the growing capacity and cost challenges in AI deployment, particularly around latency and cloud dependency. He explained how AI is approaching a point similar to Uber’s surge pricing, where real market rates for AI usage—especially inference—will soon replace current subsidized costs. This shift, possibly by 2027, will force the industry to prioritize efficiency.
Bercovici emphasized the exponential business value in increasing tokens but pointed out that costs and latency pose sustainability challenges, especially for high-stakes applications requiring accuracy and security. He described how AI agents operate in swarms, executing complex tasks with many prompt-responses, making latency a critical bottleneck that currently demands high subsidized pricing.
Reinforcement learning has emerged as a key paradigm, blending training and inference workflows to advance AI capabilities toward artificial general intelligence. On profitability, Bercovici noted there’s no one-size-fits-all infrastructure approach; organizations must tailor their strategies—on-prem, cloud, or hybrid—based on evolving needs. Ultimately, understanding and optimizing unit economics at the transaction level will be vital. The future of AI lies in smarter, more efficient large-scale application rather than reduced usage.