MCKINSEY TECH TRENDS 2026
The Hidden Chip Choice That Can Cut Your AI Cost per Answer by Up to 70%

McKinsey shows AI spending shifting to running models, where purpose-built chips can cut the cost of every answer; here is how a growing business makes sure it gets the saving.

16 min read

You will probably never buy an AI chip. But you already pay for them, every time a chatbot answers a customer, a tool drafts a proposal or a model reads an invoice. And the chip doing that work can change the price of each answer more than most owners realize.

McKinsey’s Technology Trends Outlook 2026 has a chapter on application-specific semiconductors: chips built for one kind of job rather than every job. It reads like a story for chipmakers and cloud giants. It is also a story about your AI bill, your vendor contracts and how exposed you are when supply gets tight. Here is the owner’s version.

THE 60-SECOND VERSION
The shift: AI spending is moving from building models (training) to running them (inference). Running them rewards chips designed for that one job.
The prize: In one example McKinsey cites, AWS says its inference chips deliver up to 70% lower cost per inference than comparable instances.
The momentum: Equity investment in the trend jumped from $7.5 billion in 2024 to $41.6 billion in 2025.
Your move: Ask what your AI runs on, what you pay per answer, and how easily you could switch. Most businesses cannot answer any of the three.

Butterfly chart comparing a comparable EC2 instance with an AWS Inferentia inference-chip instance. Cost per inference falls from an index of 100 to 30, up to 70 percent lower. Throughput rises from 1x to up to 2.3x.
McKinsey, Technology Trends Outlook 2026, citing Amazon Web Services: Inferentia (Inf1) instances deliver up to 2.3 times higher throughput and up to 70 percent lower cost per inference than comparable EC2 instances. Indexed and redrawn by VIVISION. “Up to” figures are vendor best cases.

PART 1 · WHY THIS IS SUDDENLY YOUR PROBLEM

The expensive part of AI is no longer building it. It is using it.

The first wave of AI spending went into training giant models on huge clusters of general-purpose graphics chips. McKinsey describes the focus now moving to inference, the work of running a trained model every time someone asks it something. Training needs raw power. Inference needs speed and efficiency, answer after answer, all day.

That changes the economics. McKinsey estimates that by 2030 inference will make up 30 to 40 percent of all data center demand and the majority of AI compute. As companies put AI into everyday workflows, each with its own speed and cost needs, data centers are mixing chip types: custom chips and specialized accelerators for specific jobs, alongside general-purpose processors.

Pie chart. Inference is estimated at 30 to 40 percent of all data center demand by 2030: the first 30 percent solid gold, the range up to 40 percent hatched, the rest dark grey for all other demand.
McKinsey estimate, Technology Trends Outlook 2026. The hatched slice shows the range of the estimate, not a measured value. Chart redrawn by VIVISION.
VIVISION insight. For a growing business, AI cost is turning from a one-off project budget into a running cost that grows with every customer, every employee and every agent you add. Running costs deserve the same discipline as rent or payroll: a unit price, a monthly review and someone accountable for it. Right now most owners we meet only see a total on a card statement.

PART 2 · WHAT OWNERS GET WRONG

Three myths about chips and your AI bill

Myth Reality
“Hardware is the cloud provider’s problem, not mine.” The provider’s hardware choices show up in your price per answer. In the AWS example McKinsey cites, the same kind of job costs up to 70% less on chips built for inference.
“One kind of chip runs everything, so the choice does not matter.” McKinsey describes data centers moving to mixed systems where different chips handle different jobs. Your vendors are already choosing. The question is whether you benefit.
“AI will only get cheaper, so I can wait.” Costs per answer are falling, but McKinsey flags memory, packaging, supply chain and energy limits that could slow AI scaling. Cheaper is likely. Smooth is not guaranteed.

The myths are VIVISION’s framing; the facts in the right-hand column come from McKinsey’s chapter.

PART 3 · FOLLOW THE MONEY

Investors added more than $34 billion in a single year

McKinsey measured equity investment in application-specific semiconductors at $7.5 billion in 2024 and $41.6 billion in 2025, and connects the surge to demand for specialized chips growing alongside AI workloads. Job postings rose 22 percent between 2024 and 2025. Start-ups focused on inference, memory and light-based chip connections are drawing money from cloud providers, venture firms and established chipmakers alike.

Waterfall chart. Equity investment was 7.5 billion dollars in 2024. New money of 34.1 billion dollars was added in 2025, for a 2025 total of 41.6 billion dollars.
McKinsey, Technology Trends Outlook 2026, trend scoring for application-specific semiconductors. The $34.1 billion increase is VIVISION’s arithmetic on McKinsey’s figures. Chart redrawn by VIVISION.
VIVISION insight. This much capital means more choice and more price competition for anyone who rents AI. It also means the big cloud platforms are using their own chips to make their offers harder to leave. McKinsey notes cloud providers increasingly use custom silicon to set themselves apart. Cheaper today can become locked in tomorrow, so price the exit before you sign.

PART 4 · WHERE DO YOU SIT?

Your lever depends on how you buy AI

PROFILE A
You use AI inside software you subscribe to
Chips are invisible to you. Your lever is the contract: usage tiers, price caps and what happens when you scale seats or volume.
PROFILE B
You call AI models through an API
You pay per use, so cost per answer is your key number. Your lever is choosing which model and which platform runs each job.
PROFILE C
You run models on cloud instances you rent
This is where the chip choice is yours. Test the same workload on general-purpose and inference-optimized options before committing.

The three profiles are a VIVISION framework, not McKinsey’s.

PART 5 · THE RISKS BEHIND THE PRICE

What could push your AI costs the other way

Concentrated supply
McKinsey notes Taiwan remains central to advanced chip manufacturing, and that geopolitics and export controls keep supply uncertain.
Memory and packaging limits
Shortages of high-bandwidth memory, advanced packaging and networking capacity could limit how fast AI capacity grows.
Power and heat
Growing AI infrastructure keeps raising pressure on power, cooling and energy efficiency.
Fast obsolescence
Chipmakers face hard trade-offs between huge factory investments and technology that ages quickly.

Paraphrased from the key uncertainties in McKinsey’s chapter. Headings are VIVISION’s.

VIVISION insight. None of these risks is yours to fix, but all of them can reach your invoice. The practical defense is portability: keep your prompts, data and workflows in a form you could move to a second provider in weeks, not months. A business with a credible plan B negotiates better and sleeps better.

QUICK ANSWERS
Will the 70% saving apply to my business?

Not automatically. It is an “up to” figure from AWS for its Inf1 instances against comparable ones, as cited by McKinsey. Real savings depend on your workload. Treat it as proof that the gap can be large, then test your own job before believing any number.

Is this only for big companies?

McKinsey says the cloud giants are the largest adopters today, and enterprise adoption is growing gradually as inference costs fall and specialized infrastructure becomes easier to access. Smaller firms benefit mostly through the platforms they rent, which is why the questions you ask vendors matter.

How mature is this trend?

McKinsey gives it an adoption score of 4 out of 5, scaling in progress, across data centers, cloud environments and enterprise workloads. It is already shaping prices, not a future bet.

Do we need to hire chip experts?

No. McKinsey finds the sharpest talent shortages in GPU and optimization skills, the core of custom-chip design, and that is the vendors’ race to run. What you need is someone who can read an AI bill, compare cost per answer and run a fair test.

YOUR MONDAY CHECKLIST
☐  List every AI tool and API you pay for, with last month’s cost.
☐  Divide cost by output for your top one: price per answer, per document or per call.
☐  Ask that vendor whether cheaper, inference-optimized options exist for your workload.
☐  Check the contract for exit terms and how you would move your data and prompts.
☐  Name one person who owns AI unit costs and reviews them monthly.

WHAT WE DO FOR CLIENTS FACING THIS

VIVISION runs an AI cost and resilience review. We map every AI service you use, work out your real cost per answer, and benchmark it against cheaper ways to run the same work, including inference-optimized options where you control the choice.

Then we tighten the contracts: usage caps, exit terms and a tested plan B provider. You leave with a lower, predictable AI running cost and a business that is not hostage to one platform’s pricing or one region’s supply chain.

Do you know what you pay for each AI answer?

Send us last month’s AI bills. We will show you your cost per answer and where it could fall.

Talk to VIVISION

Source: McKinsey & Company, “Technology Trends Outlook 2026” (Sixth edition, September 2026), Application-specific semiconductors chapter. All statistics are McKinsey’s, including the Amazon Web Services figures the report cites. Charts were redrawn by VIVISION from the published figures; the $34.1 billion increase is VIVISION’s arithmetic. The myths, the three buyer profiles, the “VIVISION insight” sections, the quick-answer opinions, the Monday checklist and “what we do for clients” are VIVISION’s own views and are not McKinsey’s.

Copyright in the original report belongs to McKinsey & Company. Cover photo: photo by Laura Ockel on Unsplash.