Reference

Viewing Usage and Costs

Monitor your workspace's AI and tool usage, costs, and execution history from the Usage page.

Monitor your workspace's AI and tool usage, costs, and execution history from the Usage page.

Before you begin

  • You must be a workspace administrator to view workspace-level usage data.

Steps

1. Open the Usage page

Navigate to the Usage page from the main navigation. The page shows a summary of your workspace's spending and activity.

2. Review the usage chart

The stacked bar chart displays monthly costs broken down by:

  • LLM Cost (cyan bars) -- token costs for AI model usage
  • Tool Cost (orange bars) -- costs for tool executions

Hover over any bar to see a detailed tooltip with the month's total cost, LLM cost, tool cost, execution count, and success/failure breakdown.

Above the chart, three summary metrics are displayed:

  • Total spend across the displayed period
  • Average monthly cost
  • Total executions

3. Check the stats cards

Four cards show key metrics at a glance:

  • Current Month -- spending so far this month, with a trend indicator comparing to last month
  • Last Month -- total spending from the previous month
  • Success Rate -- percentage of successful tool executions
  • Recent Failures -- number of failed executions in the last 24 hours

4. Review recent executions

The Recent Executions section lists the most recent tool runs. Each entry shows:

  • Status (green for succeeded, red for failed, orange for in progress)
  • Tool name
  • User who triggered it
  • Time elapsed since execution
  • Duration
  • Tool pack name

Click See more to view additional executions if available.

5. View pricing details

Click the Pricing button (blue info icon) in the page header to open the pricing dialog. This shows:

  • LLM Pricing -- input and output token costs per million tokens for each model
  • Tool Pricing -- per-call rates for tool executions

Model pricing reference

The table below lists all supported models and their token costs. These are the base costs before any workspace billing margin is applied.

OpenAI

ModelInput per 1M tokensOutput per 1M tokensSource
gpt-5.2-codex$1.75$14.00OpenAI Pricing
gpt-5$1.25$10.00OpenAI Pricing
gpt-5-mini$0.25$2.00OpenAI Pricing
gpt-5-nano$0.05$0.40OpenAI Pricing
4.1$2.00$8.00OpenAI Pricing
4o$2.50$10.00OpenAI Pricing
4o-mini$0.15$0.60OpenAI Pricing
o4-mini$1.10$4.40OpenAI Pricing
o3-mini$1.10$4.40OpenAI Pricing

Anthropic

ModelInput per 1M tokensOutput per 1M tokensSource
claude-fable-5$10.00$50.00Anthropic Pricing
claude-opus-4-6$5.00$25.00Anthropic Pricing
claude-sonnet-5$3.00$15.00Anthropic Pricing
claude-sonnet-4-6$3.00$15.00Anthropic Pricing
claude-opus-4-5$5.00$25.00Anthropic Pricing
claude-sonnet-4-5$3.00$15.00Anthropic Pricing
claude-haiku-4-5$1.00$5.00Anthropic Pricing
claude-opus-4.1$15.00$75.00Anthropic Pricing
claude-opus-4$15.00$75.00Anthropic Pricing
claude-sonnet-4$3.00$15.00Anthropic Pricing
claude-3-7-sonnet$3.00$15.00Anthropic Pricing

xAI

ModelInput per 1M tokensOutput per 1M tokensSource
grok-3$3.00$15.00xAI Pricing
grok-3-mini$0.30$0.50xAI Pricing
grok-4-1-fast-reasoning$0.20$0.50xAI Pricing
grok-4-1-fast-non-reasoning$0.20$0.50xAI Pricing
grok-code-fast-1$0.20$1.50xAI Pricing

Perplexity

ModelInput per 1M tokensOutput per 1M tokensSource
sonar$1.00$1.00Perplexity Pricing

DeepInfra-hosted Models

ModelInput per 1M tokensOutput per 1M tokensSource
deepseek-ai/DeepSeek-V4-Flash$0.10$0.20DeepInfra Pricing
deepseek-ai/DeepSeek-V4.1-Flash$0.14$0.42DeepInfra Pricing
XiaomiMiMo/MiMo-V2.6-Flash$0.14$0.28DeepInfra Pricing
XiaomiMiMo/MiMo-V2.6-Pro$0.43$0.87DeepInfra Pricing
nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B$0.50$2.50DeepInfra Pricing
zai-org/GLM-4.7-Flash$0.06$0.40DeepInfra Pricing
zai-org/GLM-5.1$1.05$3.50DeepInfra Pricing
Qwen/Qwen3.7-Max$2.50$7.50DeepInfra Pricing
Qwen/Qwen3.8-2.4T-A95B$2.00$6.00DeepInfra Pricing
moonshotai/Kimi-K2.6 (FP4)$0.75$3.50DeepInfra Pricing
zai-org/GLM-5.3 (FP4)$0.563$2.50DeepInfra Pricing

Cerebras

ModelInput per 1M tokensOutput per 1M tokensSource
zai-glm-4.7$2.25$2.75Cerebras Pricing

Self-hosted (SGLang)

Self-hosted models run on your own inference server, so there is no per-token vendor charge and token costs are recorded as zero. Tool executions are still billed at the usual rate.

ModelInput per 1M tokensOutput per 1M tokensSource
gemma-4-26b$0.00$0.00Self-hosted
qwen3.8-27b$0.00$0.00Self-hosted

Tool execution pricing

Tool executions are billed at a flat rate per call. The default global rate is $170.00 per million calls ($0.00017 per call). Individual tools or tool packs may have custom rates.

Viewing per-chat token usage

Token counts appear alongside individual chat conversations. The count shows the total tokens used in that chat session, formatted as a compact number (for example, "1.2k tokens" or "2.5M tokens").

See Troubleshooting if usage data is not loading.