Skip to main content

Costs Panel

The Costs panel provides real-time visibility into LLM spend across all governed agents. Detect runaway spending, forecast budget exhaustion, and identify optimization opportunities.

Cost Velocity Chart

The primary widget shows current burn rate (USD/minute) with a real-time sparkline:

Features

  • Updated every 5 seconds via WebSocket
  • Velocity threshold line shown as a horizontal reference
  • Configurable time windows: 1 hour, 6 hours, 24 hours
  • Alert state: When burn rate exceeds the configured threshold, the widget turns red with a pulsing border

Velocity Alert State

When a COST_VELOCITY_ALERT fires:
Shows:
  • Current rate vs configured threshold
  • Active throttle status (if action is “throttle”)
  • Which agents are contributing most to the velocity

Budget Exhaustion Forecast

A countdown widget showing when budgets will be exhausted at current spend rate:

Gauge Color Coding

Forecast Projection

A line chart showing:
  • Solid line: Actual cumulative spend over time
  • Dashed line: Linear budget allocation (expected spend)
  • Projected line: Extrapolation showing when the lines cross (exhaustion point)

Budget Scopes

Toggle between views:
  • Session: Current session budget
  • Daily: Today’s budget
  • Monthly: Monthly budget

Per-Agent Cost Attribution

A breakdown table showing which agents are responsible for spend:

Breakdown Views

  • By Agent: Individual agent costs
  • By Role: Aggregated by role
  • By Model: Which models cost the most
  • By Provider: Cost per provider (OpenAI, Anthropic, etc.)
  • By Time: Hourly/daily cost breakdown

Cost Optimization Suggestions

An actionable panel showing where money can be saved:

Categories

Context Bloat Alert

Flags requests where input tokens exceed 50% of the model’s context window:

Refresh Cadence

Optimization suggestions update daily based on the rolling 7-day usage window. This avoids performance impact from constant recomputation.

Retry Chain Costs

A table of the most expensive retry chains: Shows:
  • Top 20 most expensive retry chains
  • Whether the retry cost cap terminated the chain
  • Final outcome (did the retry succeed?)

Model Routing Log

When cost-aware model routing is configured, a log shows routing decisions: Includes:
  • Total savings from routing this month
  • Routing frequency (% of requests that hit the ceiling)
  • Warning if routing frequency > 50% (“Consider adjusting your ceiling”)

Real-Time Updates

All cost widgets update in real-time via WebSocket:
  • Burn rate: Every 5 seconds
  • Budget gauge: After each request
  • Agent attribution: Every 30 seconds
  • Optimization suggestions: Daily
  • Velocity alerts: Instant (push)