[Developers]

Token Usage Management

AI inference costs are not always intuitive. A feature that looks modest in development can become the largest cost driver in production once hundreds of analysts use it daily. The Token Usage Management module gives adm

Category: ManagementLast Updated: Apr 2, 2026
managementaireal-time

Overview#

AI inference costs are not always intuitive. A feature that looks modest in development can become the largest cost driver in production once hundreds of analysts use it daily. The Token Usage Management module gives administrators real-time visibility into exactly what is being consumed, by whom, at what cost, so model selection and quota decisions are made with current data rather than estimates.

Every AI request is tracked at the per-request level. Billable tokens apply a 1.5x multiplier to raw token counts. When provider APIs do not return exact counts, the system estimates from response length as a documented fallback.

Key Features#

  • Real-Time Token Metrics: Monitor token consumption as it happens with live dashboards showing current-day usage, hourly rates, month-to-date totals, projected month-end consumption, and quota utilisation. Break down usage by AI model, feature, user, department, or investigation for granular visibility.

  • Raw and Billable Token Tracking: Every AI request reports both raw token counts (input tokens, output tokens, total tokens) and billable token units that apply a 1.5x multiplier to raw counts. Per-request telemetry includes the AI provider, model identifier, model tier, latency in milliseconds, and a complete usage breakdown. When provider APIs do not return exact token counts, the system estimates usage from response character length as a documented fallback.

  • Cost Analysis: Detailed cost breakdowns by model, feature, user, and time period. Understand cost per request, cost per investigation, and cost per user session to make informed decisions about AI resource allocation and model selection.

  • Usage Analytics: Analyse consumption patterns with daily, weekly, and monthly trends, seasonal pattern detection, and anomaly identification. Identify top consumers, compare department usage, and track feature-level efficiency.

  • Quota Management: Set and enforce token usage quotas at the organisation, department, and individual user level. Configure monthly and daily limits with progressive alert thresholds. Choose between soft limits (warnings only) and hard limits (usage blocked) based on your governance requirements.

  • Optimisation Recommendations: Actionable suggestions for reducing token costs including model selection optimisation, prompt efficiency improvements, caching opportunities, and batch processing strategies. Each recommendation includes estimated savings and implementation complexity.

  • Predictive Forecasting: Forecast future token usage and costs based on historical patterns, seasonal trends, and growth trajectories. Assess budget risk and plan capacity to avoid unexpected cost overruns.

  • Cost Allocation: Attribute AI costs to business units, departments, projects, and investigations for chargeback and internal accounting. Track return on investment at the feature level to prioritise AI capabilities that deliver measurable value.

Use Cases#

  • Budget management with real-time cost visibility, quota enforcement, and projected spend forecasting that prevent unexpected AI cost overruns at period end.
  • Cost optimisation through model selection guidance, prompt efficiency analysis, and caching recommendations that reduce consumption without sacrificing output quality.
  • Usage governance with configurable quotas at organisation, department, and user levels to ensure fair resource allocation and prevent individual overconsumption.
  • Business intelligence through cost attribution that connects AI spending to business outcomes, enabling informed decisions about which AI features to expand or reduce.
  • Anomaly detection that identifies unusual consumption patterns, failed request surges, and unexpected cost spikes for rapid investigation.

Open Standards#

  • OAuth 2.0 and JWT Bearer Token: Token-based authentication protects typed, auditable read and write workflows across the platform.
  • JSON Web Token, RFC 7519: Every admin endpoint resolves the caller's organisation identity from an RS256-signed JWT; the organisation_id claim gates all per-tenant usage reads and quota write workflows, with JWKS-backed signature verification.
  • OAuth 2.0, RFC 6749: Administrative access is authenticated via the Bearer token pattern with scope-based permission checks, enforcing least-privilege access to usage and billing surfaces.
  • ISO 8601 / RFC 3339: All usage event timestamps, billing period boundaries, quota warning windows, and cost allocation records are serialised as UTC ISO 8601 strings, enabling unambiguous cross-system time correlation.
  • W3C Trace Context: The usage metering service accepts and persists an optional W3C trace ID alongside each integration request and authenticated workflow event, correlating token consumption records with distributed request traces.
  • OpenTelemetry (OTLP): The platform initialises the OpenTelemetry developer toolkit at start-up and exports spans via OTLP; AI usage records and quota enforcement events flow into the same observability pipeline used for service-level latency and error monitoring.
  • JSON (RFC 8259): All per-request telemetry payloads, usage metadata, quota states, cost breakdowns, and dashboard integration responses are serialised as JSON, consumed by the frontend typed integration contract client and external reporting integrations.

Getting Started#

  1. Establish Baseline: Monitor usage for 30 days to understand normal consumption patterns before setting quotas.
  2. Configure Quotas: Set organisation and department-level token limits based on your baseline data and budget.
  3. Set Up Alerts: Configure progressive alert thresholds to receive early warning as usage approaches limits.
  4. Review Recommendations: Act on optimisation suggestions starting with high-impact, low-effort improvements.
  5. Schedule Reports: Set up regular usage and cost reports for stakeholders and budget owners.

Last Reviewed: 2026-04-02 Last Updated: 2026-04-14

Ready to Build?

Get started with our APIs or contact our integration team for support.