McKinsey shows AI spending shifting to running models, where purpose-built chips can cut the cost of every answer; here is how a growing business makes sure it gets the saving.
You will probably never buy an AI chip. But you already pay for them, every time a chatbot answers a customer, a tool drafts a proposal or a model reads an invoice. And the chip doing that work can change the price of each answer more than most owners realize.
McKinsey’s Technology Trends Outlook 2026 has a chapter on application-specific semiconductors: chips built for one kind of job rather than every job. It reads like a story for chipmakers and cloud giants. It is also a story about your AI bill, your vendor contracts and how exposed you are when supply gets tight. Here is the owner’s version.

The expensive part of AI is no longer building it. It is using it.
The first wave of AI spending went into training giant models on huge clusters of general-purpose graphics chips. McKinsey describes the focus now moving to inference, the work of running a trained model every time someone asks it something. Training needs raw power. Inference needs speed and efficiency, answer after answer, all day.
That changes the economics. McKinsey estimates that by 2030 inference will make up 30 to 40 percent of all data center demand and the majority of AI compute. As companies put AI into everyday workflows, each with its own speed and cost needs, data centers are mixing chip types: custom chips and specialized accelerators for specific jobs, alongside general-purpose processors.

Three myths about chips and your AI bill
| Myth | Reality |
|---|---|
| “Hardware is the cloud provider’s problem, not mine.” | The provider’s hardware choices show up in your price per answer. In the AWS example McKinsey cites, the same kind of job costs up to 70% less on chips built for inference. |
| “One kind of chip runs everything, so the choice does not matter.” | McKinsey describes data centers moving to mixed systems where different chips handle different jobs. Your vendors are already choosing. The question is whether you benefit. |
| “AI will only get cheaper, so I can wait.” | Costs per answer are falling, but McKinsey flags memory, packaging, supply chain and energy limits that could slow AI scaling. Cheaper is likely. Smooth is not guaranteed. |
The myths are VIVISION’s framing; the facts in the right-hand column come from McKinsey’s chapter.
Investors added more than $34 billion in a single year
McKinsey measured equity investment in application-specific semiconductors at $7.5 billion in 2024 and $41.6 billion in 2025, and connects the surge to demand for specialized chips growing alongside AI workloads. Job postings rose 22 percent between 2024 and 2025. Start-ups focused on inference, memory and light-based chip connections are drawing money from cloud providers, venture firms and established chipmakers alike.

Your lever depends on how you buy AI
The three profiles are a VIVISION framework, not McKinsey’s.
What could push your AI costs the other way
Paraphrased from the key uncertainties in McKinsey’s chapter. Headings are VIVISION’s.
Will the 70% saving apply to my business?
Not automatically. It is an “up to” figure from AWS for its Inf1 instances against comparable ones, as cited by McKinsey. Real savings depend on your workload. Treat it as proof that the gap can be large, then test your own job before believing any number.
Is this only for big companies?
McKinsey says the cloud giants are the largest adopters today, and enterprise adoption is growing gradually as inference costs fall and specialized infrastructure becomes easier to access. Smaller firms benefit mostly through the platforms they rent, which is why the questions you ask vendors matter.
How mature is this trend?
McKinsey gives it an adoption score of 4 out of 5, scaling in progress, across data centers, cloud environments and enterprise workloads. It is already shaping prices, not a future bet.
Do we need to hire chip experts?
No. McKinsey finds the sharpest talent shortages in GPU and optimization skills, the core of custom-chip design, and that is the vendors’ race to run. What you need is someone who can read an AI bill, compare cost per answer and run a fair test.
VIVISION runs an AI cost and resilience review. We map every AI service you use, work out your real cost per answer, and benchmark it against cheaper ways to run the same work, including inference-optimized options where you control the choice.
Then we tighten the contracts: usage caps, exit terms and a tested plan B provider. You leave with a lower, predictable AI running cost and a business that is not hostage to one platform’s pricing or one region’s supply chain.
Do you know what you pay for each AI answer?
Send us last month’s AI bills. We will show you your cost per answer and where it could fall.
Source: McKinsey & Company, “Technology Trends Outlook 2026” (Sixth edition, September 2026), Application-specific semiconductors chapter. All statistics are McKinsey’s, including the Amazon Web Services figures the report cites. Charts were redrawn by VIVISION from the published figures; the $34.1 billion increase is VIVISION’s arithmetic. The myths, the three buyer profiles, the “VIVISION insight” sections, the quick-answer opinions, the Monday checklist and “what we do for clients” are VIVISION’s own views and are not McKinsey’s.
Copyright in the original report belongs to McKinsey & Company. Cover photo: photo by Laura Ockel on Unsplash.