Your Enterprise is Overpaying for AI and Not Because You’re Overusing It - Impetus

Your Enterprise is Overpaying for AI and Not Because You’re Overusing It

Most companies send every AI request, from “find me the leave policy” to “give me the reason for the last three quarters’ sales decline,” through the same expensive AI stack. That’s a costly way to run AI, and much of the spend is invisible on the invoice. The fix isn’t using less AI. It’s putting an intelligent layer between employees and the models.

First, what this isn’t

This isn’t an argument for rationing AI, or for asking employees to use it more sparingly. Usage going up is usually a sign the tools are working, and the last thing you want is employees second-guessing whether a question is worth asking.

Which is exactly what’s starting to happen. McKinsey’s State of AI survey for 2026 — 1,719 business leaders across 97 countries — found that one in five organizations is now limiting how much AI it uses because of what it costs to run, token costs included. In the same survey, 60 percent expect to spend more on AI over the coming year, while the share reporting any real bottom-line impact sat at 37 percent, essentially where it was a year earlier. Spending is climbing, returns are flat, and a fifth of companies have begun throttling usage to control the bill.

And the overspend isn’t caused by volume in the first place. It’s caused by every request, regardless of difficulty, being priced as though it were the hardest thing you’ll ask all year.

Don’t use a Lamborghini for a milk run

A simple request like “What’s our travel expense limit?” doesn’t need the same model, compute, or cost as “Analyse the causes of our three-quarter sales decline and recommend what we should do next.”

Yet that’s how enterprise AI usually works today. The newest and most capable frontier models handle everything — the hard reasoning tasks, yes, but also the email summaries, the policy lookups, the reformatting. All at the same premium rate.

The reason is simple: whoever built the tool wired it to a single model, chosen for the most demanding thing it might ever be asked to do. Every easy request that follows inherits that price.

Why the bill is hard to see

AI usage is billed in tokens — roughly, the chunks of text a model reads and writes. More text processed means more tokens, which means more cost. It’s a small charge per request, which is exactly why it goes unnoticed. Nobody flags a few cents. But multiply a few cents by thousands of employees making dozens of requests a day, and the number stops being small.

Worse, most organizations can’t say what kind of work that budget is going toward. They can see a total. They can’t see that a third of it went to routine lookups a far cheaper model would have handled perfectly.

Sort by pattern, not by use case

What’s needed is a layer that looks at each incoming request, works out how demanding it actually is, and sends it to the cheapest model capable of doing that job well — automatically. This isn’t asking employees to choose a tool based on how hard their question is. People shouldn’t have to make that judgement, and if you ask them to, they’ll pick the premium option every time. Which is entirely rational, and defeats the purpose.

The instinct when building this is to catalogue every way people use AI. That list never ends, and it’s out of date by the time you finish it.

The more practical move is to look for the pattern underneath. Most AI requests are one of a handful of shapes: generate something new, condense something long, extract a specific fact, or work through a multi-step task using several tools. Organizations that do this exercise typically find that six to eight patterns explain the overwhelming majority of their AI usage, and therefore their spend.

Patterns are useful because they’re stable. New use cases appear constantly; the underlying shapes don’t. And once you know the shape of a request, you can make a sensible decision about what it costs to answer.

Try the cheap option first — but check the work

Here’s the mechanism, in plain terms:

  • Classify the request — what pattern is this, how demanding is it?
  • Gateway choices — Apply best fit techniques for the pattern
  • Route it to the cheapest model that should be able to handle that pattern.
  • Check the answer against a quality standard. Correct, complete, right format?
  • Escalate if it isn’t. The request goes automatically to the stronger model, and the user gets the better answer without ever knowing a swap happened.

The whole sequence happens in the time it takes to answer a question. The person asking sees one answer, not four steps.

Think of it as triage. A scraped knee doesn’t need a surgical team, but you want one on standby and the handover to be instant when it’s needed. What you don’t want is a system that quietly hands people worse answers to save money. The quality check is what separates the two — cost savings are only real if the output still clears the bar.

There’s an honest caveat worth naming: cheap isn’t cheap if it fails often. If a particular kind of task frequently needs escalating, you’ve paid twice — once for the failed attempt, once for the retry. Which is why this only works when you measure the miss rate per pattern rather than assuming it works everywhere.

One invisible gateway instead of five separate tools

Today, AI tools tend to arrive as separate products — one chat assistant here, another built into the office suite, a third inside developer tooling. Each with its own settings, rules, and bill.

The alternative is a single gateway sitting behind the interfaces employees already use. Nothing changes for them: same tools, same questions. The gateway simply decides what happens next.

That one checkpoint is where a surprising amount of value sits:

  • Routing & escalation: Route each request based on task pattern and complexity to the model optimally priced for the required quality bar.
  • Semantic cache: Reuse answers to repeated or semantically similar questions. If five people ask the same policy question in a week, only the first needs to reach a model.
  • Policy & guardrails: Apply data-handling, security, and compliance rules consistently across models and vendors.
  • Context & entitlements across models: Preserve the conversation and the user’s permissions when a request moves from one model to another.
  • Prompt right-sizing: Remove unnecessary context and tokens before a request reaches the model, reducing cost without compromising the required output.
  • Evaluation & cost ledger: Track requests in one place so the organization can see what its AI budget is buying, broken down by team, model, and type of work.

And there is a commercial benefit. When every AI request runs through your own gateway, switching, or adding a vendor becomes a configuration change rather than a rebuild.

Build it once, not thirty times

There’s another source of waste, less obvious than token spend. When teams build AI tools independently, they build the same things: their own document summariser, meeting scheduler, report drafter — each to a slightly different standard.

A shared library of common capabilities, built once and made available through templated prompts and tools, eliminates that duplication. For a large organization this can be a bigger saving than routing itself. The cost isn’t tokens; it’s engineering time spent solving the same problem thirty times.

What you’re actually buying

Paying premium rates for routine work is a choice. Most organizations are making it by default rather than on purpose because, until recently, the infrastructure to do anything else wasn’t in place. Once it is, usage can keep growing while the cost per unit of work falls.

But cost is only half of what you get. An organization that routes its own AI traffic knows what its AI budget is actually buying, broken down by team and type of work. It can also move between vendors without rebuilding the experience employees already use. Walk into a renewal conversation with those two facts in hand and it’s a different conversation entirely. Walk in without them and you’re accepting whatever the dashboard tells you.

Frontier models are worth their price for frontier problems. The rest of the time, you’re sending a Lamborghini for the milk.

And once every AI request passes through one front door, that door becomes the natural place to enforce everything else: enterprise policies, guardrails, and AI security and trust controls, set once and applied to every tool and every team. What starts as a cost decision becomes the foundation for AI trust and maturity.

Author

Sheba Fernando,
Vice President- Engineering, Impetus Technologies

Learn more about how our work can support your enterprise