McKinsey says running AI is becoming its real operating cost. Here is how to match each job to the lightest model that does it well, before the bill grows.
Most AI bills do not blow up on day one. They creep. A pilot runs on the most powerful model available because it is easy. The pilot works, so it goes live. Usage grows, and every customer email, invoice and support ticket now goes through a model built to do almost anything, even when the job is simple.
McKinsey’s Technology Trends Outlook 2026 says AI has moved from experiments into large-scale rollout, and that building the infrastructure behind it is very expensive. The report also points to a quieter shift that matters more to a growing business: the edge is moving away from simply using the biggest model and toward running the right mix of models for each job. This article explains why AI running costs are becoming a board-level topic, and how to right-size your models before the bill does it for you.
All figures: McKinsey & Company, Technology Trends Outlook 2026, AI infrastructure and model architectures chapter, including McKinsey research and third-party sources the report cites. The 90+ GW figure is described in the report as roughly equal to California’s current electricity demand.
- The cost of running AI, every answer and every token, is becoming the operating cost of AI, and McKinsey names it as something holding back scaling.
- Behind the scenes, chips, memory, power and skilled labor are all in short supply, and data centers are taking a much bigger share of the grid.
- Small language models are becoming practical for targeted business tasks, at much lower cost, according to the report.
- VIVISION’s view: stop buying “an AI model”. Build a small portfolio of models, each matched to a job, with a cost owner and a budget.

The report contrasts two models. The version of GPT-3 that powered ChatGPT at launch had 175 billion parameters. Microsoft’s Phi-4-mini, a small language model designed for targeted uses such as on-device assistants, math and coding, has 3.8 billion. McKinsey notes that small models are showing strength in specialized tasks where accuracy, speed and cost matter more than broad reasoning, with early use concentrated in customer service and workflow automation. Size is not a price tag on its own, but it is a strong clue: a lot of everyday business work does not need the heaviest model on the market.
Your AI runs on a supply chain that is under strain
Every AI answer your business gets is produced in a data center somewhere, and McKinsey describes that whole chain as stretched. The report lists shortages of processors, high-bandwidth memory, electrical transformers and skilled labor, and says data centers can often be built faster than the power to run them arrives. It also notes that in some regions, AI-related data center demand has already contributed to notable increases in electricity prices and grid stress.

Meanwhile the money keeps pouring in. McKinsey reports equity investment in this trend of $145 billion in 2025, then nearly $384 billion in the first half of 2026 alone. At that pace, the report says, the full year would reach roughly $769 billion. Even the giants are feeling it: Alphabet raised roughly $50 billion in equity and $20 billion in debt in one quarter to help fund its build-out, per the report.

“Inference is becoming the operating cost of enterprise AI.”
Roger Roberts, partner, McKinsey, in Technology Trends Outlook 2026. Inference means a trained model producing an answer.
Myth vs reality
The biggest model is always the safest choice.
McKinsey reports that smaller models can deliver stronger domain-specific accuracy in targeted uses such as customer service and workflow automation, at significantly lower cost.
Access to the best model is the competitive edge.
The report says the advantage is shifting to orchestrating a portfolio of models that balances performance, cost, speed and governance.
Serious AI has to run in a big cloud.
Efficient models are making it more practical to run AI locally, on devices or inside your own environment, where speed, privacy, regulation or cost make the cloud less attractive, per McKinsey.
Match each job to the lightest model that does it well
This is VIVISION’s working framework, built on the portfolio idea McKinsey describes. It is a starting point for testing, not a rule.
| Type of job | Examples | Start with |
|---|---|---|
| High volume, narrow, repeatable | Sorting emails, tagging tickets, pulling fields from invoices, routine replies | Small model, tuned or prompted for the task |
| Sensitive data or needs to be fast | Client records, health or financial files, on-site or in-store tools | Small or open model run locally or in your own cloud |
| Mid complexity, uses your knowledge | Answering questions from your policies, product catalog or past projects | Mid-size model plus a search over your own documents |
| Rare, complex, high stakes | Strategy drafts, multistep analysis, complex code, negotiations prep | Frontier model, used on purpose and reviewed by a person |
From “one model for everything” to a portfolio
List every place your business uses AI, which model it calls, roughly how often, and who pays. Include tools where AI is built in and billed per seat or per use.
Place each use in the right-size map above. Circle the one with the highest volume. That is your first candidate.
Run 100 real past examples through a smaller model side by side with your current one. Score accuracy, speed and cost. Only switch if quality holds.
Give AI spend a monthly budget and a named owner, add a simple usage report, and agree the rule for when a task is allowed to use the frontier model.
Quick answers
We are not a tech company. Does AI infrastructure really affect us?
Yes, indirectly. You will not build a data center, but every AI tool you pay for runs on one. McKinsey describes inference costs as a central operating challenge for enterprises. VIVISION’s view is that the same logic applies at small scale: you pay for the model you choose, every time it runs.
Will a small model make more mistakes?
On broad, open-ended questions it may. On a narrow, well-defined task, McKinsey reports small models can deliver stronger domain-specific accuracy. That is why VIVISION always tests on your own real examples before switching anything.
Should we run AI on our own servers?
Sometimes. The report says efficient models make local or in-house deployment more practical where latency, privacy, regulation or cost make the cloud less attractive, and that open-weight models are widening those options. In VIVISION’s view it pays off mainly for high-volume or sensitive work, and only if someone can look after it.
What skills does this need?
McKinsey’s talent data for this trend shows the sharpest shortage in continuous integration and delivery skills, the know-how that moves models into live use. VIVISION’s view: you do not need a big team, but you do need one person who owns how AI is deployed and what it costs.
VIVISION runs an AI cost and model review. We map every AI tool and model your business pays for, sort each use by volume, sensitivity and difficulty, and test lighter options on your real work. You get a simple model portfolio, a budget with an owner, and a rule for when the expensive model is worth it.
If you are about to sign a larger AI contract, we help you check that its pricing matches how you will actually use it.
Not sure what your AI is really costing you?
Send us a list of the AI tools you use. We will show you where a lighter model could do the same job.
Source: McKinsey & Company, “Technology Trends Outlook 2026” (Sixth edition, September 2026), AI infrastructure and model architectures chapter. All statistics are McKinsey’s, including McKinsey research and third-party sources the report cites (among them the IEA, Stanford’s AI Index and Microsoft). Charts were redrawn by VIVISION from the published figures. The “VIVISION insight” sections, the right-size map, the 30-day plan, the quick-answer opinions and “what we do for clients” are VIVISION’s own views and are not McKinsey’s.
Copyright in the original report belongs to McKinsey & Company. Cover photo: Tyler (@tylergm) on Unsplash.