Viewing Usage and Costs
Monitor your workspace's AI and tool usage, costs, and execution history from the Usage page.
Monitor your workspace's AI and tool usage, costs, and execution history from the Usage page.
Before you begin
- You must be a workspace administrator to view workspace-level usage data.
Steps
1. Open the Usage page
Navigate to the Usage page from the main navigation. The page shows a summary of your workspace's spending and activity.
2. Review the usage chart
The stacked bar chart displays monthly costs broken down by:
- LLM Cost (cyan bars) -- token costs for AI model usage
- Tool Cost (orange bars) -- costs for tool executions
Hover over any bar to see a detailed tooltip with the month's total cost, LLM cost, tool cost, execution count, and success/failure breakdown.
Above the chart, three summary metrics are displayed:
- Total spend across the displayed period
- Average monthly cost
- Total executions
3. Check the stats cards
Four cards show key metrics at a glance:
- Current Month -- spending so far this month, with a trend indicator comparing to last month
- Last Month -- total spending from the previous month
- Success Rate -- percentage of successful tool executions
- Recent Failures -- number of failed executions in the last 24 hours
4. Review recent executions
The Recent Executions section lists the most recent tool runs. Each entry shows:
- Status (green for succeeded, red for failed, orange for in progress)
- Tool name
- User who triggered it
- Time elapsed since execution
- Duration
- Tool pack name
Click See more to view additional executions if available.
5. View pricing details
Click the Pricing button (blue info icon) in the page header to open the pricing dialog. This shows:
- LLM Pricing -- input and output token costs per million tokens for each model
- Tool Pricing -- per-call rates for tool executions
Model pricing reference
The table below lists all supported models and their token costs. These are the base costs before any workspace billing margin is applied.
OpenAI
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| gpt-5.2-codex | $1.75 | $14.00 | OpenAI Pricing |
| gpt-5 | $1.25 | $10.00 | OpenAI Pricing |
| gpt-5-mini | $0.25 | $2.00 | OpenAI Pricing |
| gpt-5-nano | $0.05 | $0.40 | OpenAI Pricing |
| 4.1 | $2.00 | $8.00 | OpenAI Pricing |
| 4o | $2.50 | $10.00 | OpenAI Pricing |
| 4o-mini | $0.15 | $0.60 | OpenAI Pricing |
| o4-mini | $1.10 | $4.40 | OpenAI Pricing |
| o3-mini | $1.10 | $4.40 | OpenAI Pricing |
Anthropic
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| claude-fable-5 | $10.00 | $50.00 | Anthropic Pricing |
| claude-opus-4-6 | $5.00 | $25.00 | Anthropic Pricing |
| claude-sonnet-5 | $3.00 | $15.00 | Anthropic Pricing |
| claude-sonnet-4-6 | $3.00 | $15.00 | Anthropic Pricing |
| claude-opus-4-5 | $5.00 | $25.00 | Anthropic Pricing |
| claude-sonnet-4-5 | $3.00 | $15.00 | Anthropic Pricing |
| claude-haiku-4-5 | $1.00 | $5.00 | Anthropic Pricing |
| claude-opus-4.1 | $15.00 | $75.00 | Anthropic Pricing |
| claude-opus-4 | $15.00 | $75.00 | Anthropic Pricing |
| claude-sonnet-4 | $3.00 | $15.00 | Anthropic Pricing |
| claude-3-7-sonnet | $3.00 | $15.00 | Anthropic Pricing |
xAI
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| grok-3 | $3.00 | $15.00 | xAI Pricing |
| grok-3-mini | $0.30 | $0.50 | xAI Pricing |
| grok-4-1-fast-reasoning | $0.20 | $0.50 | xAI Pricing |
| grok-4-1-fast-non-reasoning | $0.20 | $0.50 | xAI Pricing |
| grok-code-fast-1 | $0.20 | $1.50 | xAI Pricing |
Perplexity
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| sonar | $1.00 | $1.00 | Perplexity Pricing |
DeepInfra-hosted Models
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| deepseek-ai/DeepSeek-V4-Flash | $0.10 | $0.20 | DeepInfra Pricing |
| deepseek-ai/DeepSeek-V4.1-Flash | $0.14 | $0.42 | DeepInfra Pricing |
| XiaomiMiMo/MiMo-V2.6-Flash | $0.14 | $0.28 | DeepInfra Pricing |
| XiaomiMiMo/MiMo-V2.6-Pro | $0.43 | $0.87 | DeepInfra Pricing |
| nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B | $0.50 | $2.50 | DeepInfra Pricing |
| zai-org/GLM-4.7-Flash | $0.06 | $0.40 | DeepInfra Pricing |
| zai-org/GLM-5.1 | $1.05 | $3.50 | DeepInfra Pricing |
| Qwen/Qwen3.7-Max | $2.50 | $7.50 | DeepInfra Pricing |
| Qwen/Qwen3.8-2.4T-A95B | $2.00 | $6.00 | DeepInfra Pricing |
| moonshotai/Kimi-K2.6 (FP4) | $0.75 | $3.50 | DeepInfra Pricing |
| zai-org/GLM-5.3 (FP4) | $0.563 | $2.50 | DeepInfra Pricing |
Cerebras
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| zai-glm-4.7 | $2.25 | $2.75 | Cerebras Pricing |
Self-hosted (SGLang)
Self-hosted models run on your own inference server, so there is no per-token vendor charge and token costs are recorded as zero. Tool executions are still billed at the usual rate.
| Model | Input per 1M tokens | Output per 1M tokens | Source |
|---|---|---|---|
| gemma-4-26b | $0.00 | $0.00 | Self-hosted |
| qwen3.8-27b | $0.00 | $0.00 | Self-hosted |
Tool execution pricing
Tool executions are billed at a flat rate per call. The default global rate is $170.00 per million calls ($0.00017 per call). Individual tools or tool packs may have custom rates.
Viewing per-chat token usage
Token counts appear alongside individual chat conversations. The count shows the total tokens used in that chat session, formatted as a compact number (for example, "1.2k tokens" or "2.5M tokens").
See Troubleshooting if usage data is not loading.