Built for whoever owns the AI spend line

Cost reduction on the side of governance.

If you're the one explaining the LLM bill to finance: cache and body-shrink cut token spend on repeat traffic today, and named routing presets let you point a call at a cost tier you choose. Automatically picking the cheapest model for a given call, sometimes called model arbitrage, is an idea we have designed toward and have not shipped. Cost is a side effect of the design, not the wedge: governance is why teams install it in the first place.

Coming soon

The cost page collects the practitioner-facing levers as they ship. The principle: governance is the wedge; cost is the customer's first quick-win, and this page says exactly where the real-vs-roadmap line is.

  • Cache + body-shrink: reuse answers to near-identical requests and trim redundant prompt bytes; hit rates vary by workload (shipped, gateway-routed traffic)
  • Named routing presets (auto, eco, premium, free, reasoning): configure which tier a call uses per environment (shipped as configuration)
  • Automatic complexity-tier routing (model arbitrage): picking the cheapest viable model per call automatically; an idea, not built, no cost-routing decision ships today
  • Context-window auto-swap: automatically move to a larger context window on overflow; concept stage, not built
  • Per-tenant cost visibility: spend measurement on gateway-routed traffic; automatic budget-ceiling refusal is on the roadmap, not shipped
Talk to us