Diagram showing how an agent broker connects clients with different service providers
Agent Fabric for Cost Management

View and control what your AI actually costs

Price every request against the rate you pay, cut the token waste agents never needed, and cap spend before it overruns, all in a single plane.

Nail down your true token spend.

View and govern your AI spend per team, application, and use case, to keep your AI budgets in check and automatically stop run away agent loops.

Cut waste and optimize usage.

Most token spend is: full of schemas and bloated responses shipped on every request. Gateway policies strip the waste without touching your servers or client code.

Set and enforce your budget.

Budgets are enforced automatically, per instance, or per business unit, so enforcement kicks in before the overrun, not after the invoice explains it.

Screenshot of Model Proxy pricing each request against the provider rate

Start with a cost figure you can trust

When AI traffic runs through a proxy, everything needed for a real cost figure is captured together: the model that responded, the tokens it burned, and the rate your provider charges. With Agent Fabric, every request is priced against the rate you actually pay, so you can see exactly what each agent costs. The figure on the dashboard is the figure on the invoice.

Diagram of MCP token optimization policies reducing payload size

Stop paying for tokens your agents never needed

Track which MCP servers and agents are causing the most token waste, and implement cost optimization through Agent Fabric that controls your spend, without slowing your agents down.

Screenshot of Model Wallet enforcing a live budget in the Unified Dashboard

Right-size the spend that’s left, then hold the line

Recommendations catch the waste; and Model Wallet in Agent Fabric gives live budget enforcement for proactive, locked-in spending controls your agents can’t overrun.

Cost Management Frequently Asked Questions

A set of capabilities built into Agent Fabric that price, attribute, optimize, and cap AI spend across agents, MCP servers, and LLMs, from the gateway your traffic already runs through.

Model Proxy is generally available now. MCP tool optimizations, Cost Optimization for agents and model proxies, Model Wallet, and the Unified Dashboard arrive later this quarter.

Model Proxy sits in the traffic path and prices each request against the rate your provider actually charges, so dashboard spend reconciles against the invoice. For models outside the proxy, MuleSoft reports verifiable token volume and call counts.

No. The optimization policies run at the gateway. Progressive Disclosure, Smart Response Trimming, and TOON compression cut token waste without modifying your servers or client code, and the decision to apply any change stays with you.

Model Wallet enforces budgets per agent, per instance, or per business unit, tracked live as spend accrues. When a wallet runs down, enforcement kicks in before the overrun rather than after the invoice, and forecasting flags a projected overrun days ahead.

+

Esta página está disponible en español

Ver en español