Diagram showing how an agent broker connects clients with different service providers
Agent Fabric for Cost Management

Control what your AI costs.

See what every AI request actually costs, cut the waste agents don't need, and cap spend before it runs over — all from one place

Nail down your true token spend.

Track spend by team, app, and use case in real time, and automatically shut down runaway agent loops before they blow the budget.

Cut waste and optimize usage.

Gateway policies automatically strip oversized schemas and bloated responses from every request — no server or client changes required.

Set and enforce your budget.

Cap spend before it overruns. Set automatic spend ceilings per agent or business unit, so budgets are enforced before you ever see an overage on the invoice.

Screenshot of Model Proxy pricing each request against the provider rate

Start with a cost figure you can trust

When AI traffic runs through a proxy, everything needed for a real cost figure is captured together: the model that responded, the tokens it burned, and the rate your provider charges. With Agent Fabric, every request is priced against the rate you actually pay, so you can see exactly what each agent costs. The figure on the dashboard is the figure on the invoice.

Diagram of MCP token optimization policies reducing payload size

Stop paying for tokens your agents never needed

Track which MCP servers and agents are causing the most token waste, and implement cost optimization through Agent Fabric that controls your spend, without slowing your agents down.

Screenshot of Model Wallet enforcing a live budget in the Unified Dashboard

Right-size the spend that’s left, then hold the line

Recommendations catch the waste; and Model Wallet in Agent Fabric gives live budget enforcement for proactive, locked-in spending controls your agents can’t overrun.

Get more information and take back control today

Agent Fabric for Governance
Agent Fabric for Governance

Explore MuleSoft’s AI & API Gateway.

See how one governance layer secures and manages every API, agent, and model across the gateways you already run.

Cost management
Cost management

Read the cost management deep dive.

Walk through Model Proxy, MCP optimization, and Model Wallet, and how each turns AI spend into a managed line item.

Agent Fabric Demo
Agent Fabric Demo

See it in action.

Watch cost attribution, optimization, and budget enforcement work end-to-end on live AI traffic.

Cost Management Frequently Asked Questions

A set of capabilities built into Agent Fabric that price, attribute, optimize, and cap AI spend across agents, MCP servers, and LLMs, from the gateway your traffic already runs through.

Model Proxy is generally available now. MCP tool optimizations, Cost Optimization for agents and model proxies, Model Wallet, and the Unified Dashboard arrive later this quarter.

Model Proxy sits in the traffic path and prices each request against the rate your provider actually charges, so dashboard spend reconciles against the invoice. For models outside the proxy, MuleSoft reports verifiable token volume and call counts.

No. The optimization policies run at the gateway. Progressive Disclosure, Smart Response Trimming, and TOON compression cut token waste without modifying your servers or client code, and the decision to apply any change stays with you.

Model Wallet enforces budgets per agent, per instance, or per business unit, tracked live as spend accrues. When a wallet runs down, enforcement kicks in before the overrun rather than after the invoice, and forecasting flags a projected overrun days ahead.

+

Esta página está disponible en español

Ver en español