Fueled by an "exponential increase" in generative AI usage, companies like Meta, Uber, and Walmart are actively seeking ways to reduce their AI spending. While some companies, like Meta, are still on track to spend billions on AI this year, they are now focused on finding cost efficiencies to achieve similar or better business results. The shift comes after a period where employees, engaging in what they called "tokenmaxxing," often used the most powerful and expensive AI models for all tasks, leading to unexpected and substantial bills.
Tokens, which are fundamental units of data processed by AI models (often akin to word fragments), are the primary drivers of these costs. Google CEO Sundar Pichai noted that Google processes about 3.2 quadrillion tokens monthly. The cost-efficiency challenges are leading companies to adopt various strategies. For instance, Google suggests re-routing AI work to cheaper models like Gemini 3.5 Flash, which offers "frontier-level capabilities at less than half the price of comparable frontier models." AT&T's chief AI officer, Andy Markus, stated that engineers are using less advanced, and thus cheaper, AI models for everyday tasks, reserving the most powerful models for complex tasks, which can save up to 90%.
Beyond model selection, companies are emphasizing prompt efficiency to reduce token consumption. ManpowerGroup, for example, reduced the average number of follow-up questions needed for internal labor-market tool queries from 10 to 4 through more efficient prompting. Some companies are also tracking "agentic work units" instead of just tokens to measure output and value. Additionally, the use of "forward-deployed engineers" (FDEs) is gaining traction to design systems with cost requirements in mind, including selecting appropriate models to manage per-token costs. DevRev is implementing a memory layer between AI agents and data sources to cut token load and improve data movement efficiency.
The realization of these significant costs has prompted a shift from promoting high AI usage to implementing limits. For example, Uber quickly exceeded its projected AI spending for the year within four months and has since placed monthly limits on AI coding tools. Pylon CEO Marty Kausas faced a potential $1.4 million bill and decided to set token ceilings for some non-technical employees, signaling an end to unlimited spending. This emphasis on controlling costs is also driving a move towards outcome-based pricing, where the value unit is based on results rather than fragmented words, as envisioned by Seth, a senior director analyst at Gartner.