RC
The unit price of AI has been falling faster than almost any input a go-to-market team buys. LLM inference costs dropped more than 90% in about two years, per Stanford HAI. So here is the awkward question for anyone who owns an AI budget: why is the bill going up?
Gartner has a name for it. They call it the inference paradox, and they project that AI inference costs per agentic workflow will rise more than fivefold through 2028. Per-token prices keep falling. Multistep agents consume more tokens, and often more expensive ones, so the total climbs anyway.
Most teams will respond by shopping for a cheaper model. That is the wrong lever. The price of a token is not the problem. The number of times you pay for the same context is.
Where the tokens actually go
A chat assistant answers a question once. An agent plans, retrieves, reasons, calls a tool, checks the result, and tries again. Every one of those steps is a model call, and most of them carry the full context along for the ride. One "task" can be dozens of calls.
Now put that inside a GTM stack. Take a single account in a live deal, and count who needs to understand it:
The research agent pulls the CRM record, recent news, and the last few calls to write an account brief.
The outbound agent pulls the CRM record and the email history to draft a follow-up.
The call-prep agent pulls the CRM record, the call transcripts, and the open opportunity to build an agenda.
The forecasting agent pulls the opportunity, the activity log, and the same transcripts to score the deal.
Four agents. The same account. The same systems, queried separately, summarized separately, and paid for separately. Nothing any of them learned is available to the next one, so tomorrow they all do it again.
You are not buying four units of intelligence. You are buying one unit of context four times, and getting four slightly different versions of the truth as a bonus.
The multiplier nobody budgeted for
This would be a rounding error if teams ran one agent. They don't. KPMG's Q2 2026 Pulse found that organizations orchestrating multiple agents across workflows jumped from 9% to 18% in a single quarter.
In a stack where every agent wires its own integrations and its own retrieval, cost does not scale with value. It scales with agent count times workflow depth. Add a fifth agent and you don't add one more line of spend. You add another full copy of the retrieval work the first four are already doing.
At pilot scale this is invisible. Ten reps, one agent, a few hundred runs a week. It shows up at rollout, when the finance team asks why the inference line tripled while pipeline stayed flat.
Why a cheaper model won't save you
Model routing helps. Sending simple tasks to a smaller model is sensible hygiene, and you should do it.
But it optimizes the unit price of a call, and the unit price is the part of the equation that is already falling on its own. The part that is rising is volume, and volume is set by architecture: how many times the stack re-fetches, re-reads, and re-summarizes what it already knew.
We have seen this pattern before. Cloud compute got cheaper every year, and cloud bills still grew. The teams that got spend under control did not do it by waiting for the next price cut. They did it by fixing how the workload was built. Agentic AI is on the same path, just faster.
What actually bends the curve
Three design choices decide whether your agent bill grows with value or with headcount of agents.
Retrieve once, share everywhere. Account context should be assembled once, kept current, and read by every agent that needs it. An agent that starts from a structured, shared picture of the account needs a far smaller context window than one that starts from a pile of raw documents.
Connect once. If five agents each maintain their own connection to the CRM, the call recorder, and the document store, you pay five times in tokens and five times in maintenance. One governed connection that every agent uses removes the duplication at the source.
Measure cost against outcomes, not tokens. A token dashboard tells you what you spent. It does not tell you whether it was worth it. The useful number is cost per deal stage advanced, per meeting booked, per risk caught. Without that, every agent looks equally justified right up until renewal.
This is the reason we built wysdym as a layer and not as one more agent. wysdym is the operating layer for agentic GTM. It grounds every agent in your GTM truth, runs the motion under your governance, and gets sharper with every deal you close. In practice that means Cortex holds the shared intelligence and memory every agent reads from, the Gateway is one MCP door to your stack, and Observe is designed to tie agent activity to deal outcomes. The platform page walks through how the pieces fit.
It is also why we don't resell tokens. Your model bill stays on your own account, so we have no reason to want it bigger.
To be clear about what this is: an architectural argument, not a benchmark. We are building with design partners and have no production cost numbers to publish. We won't invent any.
Four questions to ask before the next renewal
You don't need new tooling to find out how exposed you are. Ask:
How many of our agents pull the same account, contact, or deal context from the same systems?
Can anyone show the inference bill broken down by agent and by workflow?
For our most expensive agent, what is the cost per outcome, in a unit sales would recognize?
If we double the number of agents next year, what happens to the bill? Does anything get cheaper because it is shared?
If the honest answer to the last one is "it doubles," the paradox is already in your stack. Cheaper tokens will not get you out of it. A shared foundation will.
More agents won't grow revenue. Compounding intelligence will. If you are wrestling with this in your own stack, we are working through it with a small group of design partners, and we would like to compare notes.
