THE ERA OF PICKING JUST ONE AI LAB IS OVER, SAYS VERCEL CEO
Companies are moving from single-model commitments to mixing OpenAI, Anthropic, Gemini, and Chinese labs as the token-burning era ends.
by editor4 min readcomments soon

Vercel CEO Guillermo Rauch has a short message for every startup that bet its entire AI stack on a single model provider: that strategy is done. Companies are now splitting their AI workload across multiple labs, routing prompts to whichever model delivers the best price-performance tradeoff for a given task, the way they once spread workloads across multiple cloud providers like AWS and Azure.
"Last year, there were a lot of people picking one lab partner — saying they would build everything on OpenAI or Anthropic," Rauch said. "You can use OpenAI, you can use Anthropic, or you can use Gemini," "every piece is plug and play."
The shift is not theoretical. Coinbase CEO Brian Armstrong said on X in June that he has been experimenting with Chinese models including Z.ai's GLM-5.2 and Kimi AI's K2.7 as defaults for his engineers, and deploying model routing systems that automatically send prompts to the most appropriate model for the task at hand. Armstrong is not an outlier. Rauch made his comments at the HumanX Conference as companies across tech are confronting a hard truth: the money flowing into AI has not translated into proportional value for customers.
THE END OF TOKEN BURNING
The instruction to employees six or twelve months ago was to push as many tokens through as possible. Experiment fast. Ship second, ask questions later. That phase is over. Companies are now actively looking for ways to cut AI spending or use the models more efficiently. Rauch described the shift in terms that will be familiar to anyone who has watched a technology go through the hype cycle. The initial phase, he said, was "all about prototyping", but now the industry is "getting into the realities of agents in production, and some of the challenges."
The math changes when you stop prototyping and start serving real traffic. A cheap model that handles 80% of your use cases well and costs a tenth of what GPT-4 charges for inference is not a compromise. It is a business necessity. Rauch noted that some models offer "awesome price/performance characteristics", a signal that the ecosystem has matured enough to compete on cost.
THE MULTI-CLOUD PARALLEL
The most useful framework for understanding what is happening is the multicloud transition of the 2010s. Companies that once committed exclusively to Amazon Web Services or Microsoft Azure eventually realised vendor lock-in was expensive and limiting. They adopted a multi-cloud strategy, picking the best storage, compute, and database services from different providers. Rauch sees the same pattern forming in AI. The model layer is becoming a stack, not a single dependency. A company might use Gemini for long-context reasoning, Anthropic for instruction-following, a fine-tuned open model for latency-sensitive inference, and a Chinese lab's model for certain cost-constrained tasks. The interfaces are standardising enough to make that work without a rewrite every time a provider changes its pricing or discontinues a model.
THE VALUE QUESTION
The underlying driver is not architectural elegance. It is the growing realisation that the AI industry's revenue growth has not been matched by customer outcomes. Companies paid for API access, premium tiers, and compute credits, then found that the output quality did not move their product metrics enough to justify the spend. Rauch is in a position to see this pattern across hundreds of companies. Vercel is a San Francisco-based cloud platform that helps developers host and launch websites and applications, which means its customers are building and deploying software that increasingly embeds AI features. When those customers hit the ceiling on value per token burned, they start shopping for alternatives.
WHAT COMES NEXT
The winners in this new phase will be the model providers that optimise for production economics rather than benchmark bragging rights. A model that is cheap enough to run at scale, fast enough for user-facing latency, and good enough for the task will beat a model that scores slightly higher on a leaderboard but costs five times as much per query. The losers will be the companies that built their pitch around exclusivity contracts and the assumption that once a developer committed to one API, the switching cost would keep them locked in. That assumption is collapsing, and Rauch is effectively telling the industry that the window for lock-in has closed.
For the companies building on these models, the math is straightforward. If you have five tasks and five models, each optimised for one of them, your average cost per task drops while your reliability per task rises. That is the argument winning boardroom debates right now. And as Rauch put it, the era of betting everything on a single lab partner is not evolving. It is over.
what did you make of it?
more from ai
ai
TSMC ADDS $100 BILLION TO ARIZONA CHIP BET, TOTAL HITS $265 BILLION
The additional investment will build at least four more 2nm fabs and advanced packaging, bringing the company's total US commitment to $265 billion.
ai
META WILL ALERT PARENTS IF TEENS DISCUSS SUICIDE WITH META AI
The opt-in feature flags self-harm references in chatbot conversations, with human review before any notification is sent.
ai
ROBLOX'S "BUILD" LETS ANYONE MAKE A GAME FROM THEIR PHONE WITH AI
The new toolset, launching July 28, turns text prompts into playable experiences and puts game creation on iPhone and iPad.
ai
ZOOX REALLS ENTURE ROBOTAXI FLEET OVER SMOKE DETECTION FAILURE
A robotaxi drove into an active fire scene obscured by smoke. NHTSA called emergency scenes not edge cases.





