Strategy & AI in business

When enterprise AI starts costing more than humans: the economic turning point of May 2026

Post written by Lorenc, Wiven AI Team · June 2026 · 7 min read

Microsoft removes Claude Code from its own engineers. Uber exhausts its 2026 AI budget by April. The "usage-based API" model meets accounting realities.

And with it, a fundamental question: can we still afford an AI stack when we don't control either the cost or the supplier?

1. The event

In May 2026, two weak signals became major warnings.

Microsoft restricted its own developers' access to Claude Code — a code assistance tool they used extensively internally. The reason? Costs per developer had reached between between 500 and 2,000 dollars per month, This is more than some part-time human positions. Management decided to limit API calls and favor its own models, which are less efficient but more controlled.

At the same time, Uber revealed that it had consumed its entire annual AI budget in four months only. The usage-based billing — tokens consumed, API calls, computation time — had exploded far beyond forecasts. Result: a temporary freeze on new AI projects and a widespread internal audit.

«We didn’t have a cost problem. We had a visibility problem. Nobody knew what each team was actually spending until the budget was gone. »

— Sundeep Gupta, VP Engineering, Uber (May 2026)

These two cases are not anecdotes. They are the first visible signs of a structural reversal AI, which was supposed to reduce costs, is actually causing them to explode — invisibly, without limits and often irreversibly.

Key figures

500 – 2000 $
Monthly cost per developer of Claude Code at Microsoft
4 months
Time required for Uber to exhaust its annual AI budget by 2026
June 30, 2026
Microsoft's restriction deadline for third-party AI tools
June 1, 2026
The price increase for the OpenAI GPT-4.1 API has now taken effect.

2. The mechanism: why usage-based pricing is exploding

The dominant business model for enterprise AI today is based on a simple principle: You pay for what you consume.. Input tokens, output tokens, API calls, GPU time.

In theory, it's flexible. In practice, it's a trap.

Three factors converge to cause budgets to drift:

Price escalation

The models are becoming more powerful—and more expensive. Each new version of GPT, Claude, or Gemini is more performant, but also more expensive per token. OpenAI has increased its API fees from 15 to 30 % with each generation since 2024.

Unpredictable use

An AI agent processing invoices doesn't consume the same amount of resources depending on whether it receives 50 or 500 documents per day. Yet no one bases an AI budget on peak workload. Teams only discover the bill at the end of the month.

Shadow AI

In most large organizations, teams use AI tools without centralized validation. Each department subscribes to its own APIs, creates its own agents, and generates its own costs—without any visibility for the IT department. This phenomenon, called shadow AI, represents between 30 and 60 % of actual AI spending.

The result

A budget of CHF 50,000 ended up at CHF 180,000. And nobody saw the overspending coming.

3. Dual Dependency

The problem of costs is only the visible part of a deeper issue: the dependence.

Technology dependence First, when a company builds its entire AI infrastructure around a single provider—OpenAI, Anthropic, Google—it ties its operations to that provider's decisions. Price changes, modifications to terms of service, performance degradation, model removal: the company suffers without recourse.

This is exactly what happened to Microsoft: when Anthropic revised its commercial licensing terms for Claude Code, Microsoft was faced with a binary choice—pay more or cut off access. There was no plan B.

Geographical dependence Next. Almost all consumer AI APIs pass through data centers located in the United States. This means that your customer data, your internal documents, your business processes pass through infrastructures subject to the US CLOUD Act.

For a Swiss SME, this represents a double risk: legal (potential non-compliance with the nLPD and the GDPR) and operational (Service may be interrupted in case of sanctions or regional access restrictions).

The key point

L'’agnosticism about models It is not a technical luxury. It is a condition for operational survival.

4. The sovereign and multi-model alternative

In light of these observations, an alternative model is emerging. It is based neither on a single supplier nor on opaque billing.

It rests on three pillars:

Sovereignty of deployment. AI is deployed directly on the client's infrastructure — on-premise or on a Swiss cloud (Exoscale). The data remains within the defined scope.

Agnosticism on models. The architecture isn't tied to a single model. If GPT-4.1 becomes too expensive, we switch to Mistral, LLaMA, or an open-source model. That's the approach. multi-models : the value is in the orchestration, not in the engine.

Predictable, contractually agreed and capped cost. No token-based billing. No surprises at the end of the month. A clear subscription that covers deployment, maintenance, and upgrades.

Forward Deployed AI

This is the approach that Wiven has applied since its creation: AI agents deployed on site, integrated with existing tools (SAP, Abacus, Odoo), managed with the client — not a generic SaaS imposed from the outside.

5. Five questions to ask before signing an AI contract in 2026

These questions are not neutral. They reflect what we believe to be the defining criteria for a successful AI project. We are sharing them transparently so that every company can evaluate its options—including those outside of Wiven.

01

Billing template

Does my provider bill me based on usage (tokens, API calls, GPU) or on a predictable model? If it's based on usage: is there a contractual limit? What happens if my usage doubles in three months?

02

Data location

Where does my data pass through? Which datacenter, in which country, under which jurisdiction? Do I have a contractual guarantee that my data does not leave Switzerland?

03

Supplier dependency

What happens if my AI model provider raises their prices by 30 %? Changes their terms? Removes the model I'm using? Do I already have a technical alternative in place?

04

Ownership and portability

Do I own my AI agent, its prompts, and its training data? Can I migrate to another provider without starting from scratch?

05

Transparency of actual costs

What is the true cost of each deployed AI agent, including hidden costs (training, integration, maintenance, model evolution)? Does my provider give me this visibility?

The end of a certain naiveté

Enterprise AI is entering a new phase. One where productivity promises clash with accounting realities. One where CIOs discover their AI budget has been exhausted before the end of the first half of the year.

This turning point is not bad news. It's a opportunity to regain control.

The companies that will emerge victorious from this period are those that have made three clear choices: a cost model predictable, an architecture multi-models, and a deployment sovereign.

Not because it's trendy. Because it's the only way to build AI that lasts — without depending on a supplier, without being subject to their prices, and without compromising data security.

Sources

  1. Business Insider, «Microsoft restricts employee use of Claude Code over cost concerns,» May 2026.
  2. The Information, «Uber burned through its 2026 AI budget by April», May 2026.
  3. OpenAI, «API Pricing Updates — GPT-4.1», official announcement, May 2026.
  4. Gartner, «Predicts 2026: AI Cost Management Will Become a Board-Level Priority,» Q1 2026 Report.
  5. BCG AI Radar, «The Hidden Cost of Shadow AI in Enterprise,» April 2026.
  6. Federal Data Protection Commissioner (FDPIC), «Recommendations on cross-border data processing», updated 2026.

To discuss your AI strategy, your actual costs, or your vendor dependency — let's talk.