build-in-public

build-in-public

Bring Your Own Model. We Don't Resell Tokens.

Bring Your Own Model. We Don't Resell Tokens.

Bring Your Own Model. We Don't Resell Tokens.

RC

Rob Catalano - Co-founder -

Rob Catalano - Co-founder -

-

-

3 min read

3 min read

Every vendor selling AI to a go-to-market team has made one decision you will never see on the demo: whether the model bill runs through them, or through you.

It sounds like a procurement detail. It isn't. It's the single line in the architecture that determines whether your vendor makes more money when your agents get more efficient, or less. We made our call early, and it's worth explaining in public — because the reasoning is more useful to you than the answer.

The incentive nobody puts on a slide

Enterprises spent $37 billion on generative AI in 2025, 3.2x the year before, with the largest slice — $19B — landing at the application layer, per Menlo Ventures. Worldwide AI spending is forecast to hit $2.59 trillion in 2026, up 47% in a single year, per Gartner. A lot of that money now moves through vendors who resell you inference at a markup.

When your vendor owns the model relationship, three things follow. Your bill is opaque — you get a seat price or a credit pack, not a token line. You can't switch models when a cheaper or better one ships, which in this market is roughly monthly. And every inefficiency in the vendor's retrieval — every redundant lookup, every over-stuffed context window — shows up as revenue for them and cost for you.

That last one is the real problem. LLM inference costs dropped more than 90% in about two years, per Stanford HAI. If your bill didn't fall with them, someone kept the difference.

What we did instead

wysdym is the operating layer for agentic go-to-market. It isn't an agent. You bring whichever agents you want — Claude, OpenAI, LangGraph, something your team built — and wysdym is the layer underneath them: shared grounding, governed action, outcome feedback.

That architecture makes the billing decision easy. The model keys stay yours. Inference runs on your own account, at your negotiated rate, against whichever model you picked this quarter. We charge for the operating layer — the graph, the skills, the governance, the observability — not for the tokens that pass over it.

Three consequences, in order of how much they matter to you:

You can switch models without switching vendors. Model prices and capabilities have moved faster than any procurement cycle can track. If your operating layer is model-agnostic, that volatility is an advantage. If your vendor resells you one model, it's a lock-in.

You see the actual bill. Your cloud provider invoices you for inference. We invoice you for the layer. Nobody has to reconcile a credit pack against what your agents actually did.

Our incentive points the same direction as yours. This is the part that matters. In wysdym, your stack connects once through the Gateway — one MCP door to Salesforce, HubSpot, Gong, Slack, Drive, Notion — and every agent reads from the same per-tenant knowledge graph. That's Cortex, and it's in production today. When five agents need the same account context, they don't each go re-retrieve it and re-pay for it. They share the grounding.

If we resold you tokens, that de-duplication would be a revenue cut. Because we don't, it's just a better product.

The trade we accepted

Being honest about the downside: reselling inference is a genuinely good business. It grows with usage automatically, it's simple to price, and it lets a vendor show a bigger contract value on day one. Giving that up means our revenue has to come from the layer being worth paying for on its own merits — the shared memory, the typed skills, the write-approval queues and per-agent permissions, the attribution of agent actions to deal stages.

It also means we have to be useful to a team that already has strong opinions about its model stack, rather than replacing those opinions. That's a harder sale and a better product. We think it's the right side of the trade for a category that's still forming — but it is a trade, and I'd rather say so than pretend it was free.

More agents won't grow revenue. Compounding intelligence will — and it compounds faster when nobody in the stack is paid to make it slower.

If you're running more than one GTM agent and starting to feel the token bill, we should talk. We're taking on a small number of design partners now.

Enjoyed the read? There’s no book on agentic GTM.

Enjoyed the read? There’s no book on agentic GTM.

Enjoyed the read? There’s no book on agentic GTM.

so we’re writing it weekly

so we’re writing it weekly

“pearls of wysdym”, weekly, 5 min read, free