Uber and Stripe Ditch Premium AI Models. Smart Routing Saves the Bill. The Model Was Never the Point.
According to The Pragmatic Engineer, companies including Uber, Pinterest, Stripe, Coinbase, Ramp, and AT&T are reporting significant cost savings by dropping proprietary AI models and adopting smart model routing instead. The piece also highlights a new CPU shortage emerging after the GPU and memory shortages driven by AI demand. AI agents that use tool calls consume substantially more CPU than traditional inference workloads.
The principle here is route optimization over model worship. Companies are learning that not every task needs the most expensive brain. The mechanism is intelligent routing. You send simple queries to cheap models and reserve heavy compute for hard problems. The savings are real. The CPU shortage is the second-order effect nobody planned for. Agents that call tools burn general compute, not just GPU cycles. Predicting which hardware you will need is now a supply chain problem, not a preference.
The Pragmatic Engineer reports that Uber, Pinterest, Stripe, Coinbase, Ramp, and AT&T are leading this shift. These are companies with enough engineering talent to build internal routing logic rather than paying premium per-query pricing to a single vendor.
- Open a free account on a model routing service like OpenRouter or LiteLLM and connect it to a consumer chatbot interface.
- Set up routing rules that send simple questions to a cheap or open model like Llama or Haiku, and route complex reasoning tasks to a premium model.
- Send ten mixed questions through the interface and observe which model each query triggers. You have just replicated the cost optimization strategy that Uber and Stripe deployed at enterprise scale. The principle is identical. Only the budget differs.