Token Counting
Eneo tracks token usage for every LLM request to provide accurate insights into AI consumption across the platform.
How It Works
Every request to a language model records two values:
- Input tokens — the prompt, system instructions, conversation history, and any tool definitions sent to the model
- Output tokens — the response generated by the model, including any internal reasoning
These counts are aggregated and displayed in the Insights and Admin dashboards.
Accurate Counting With Major Providers
When using hosted providers such as OpenAI, Anthropic, or Azure, Eneo receives the actual token counts directly from the provider’s API response. This is the most accurate method because:
- The provider uses its own tokenizer, purpose-built for its models
- All overhead is included — system prompt formatting, tool schemas, and internal message framing
- Reasoning tokens (used by models like OpenAI o-series or Claude with extended thinking) are captured
This works for both streaming and non-streaming requests, and token usage is correctly accumulated even when the model makes multiple rounds of tool calls.
Estimation for Self-Hosted Models
Self-hosted models (such as those running on vLLM or Ollama) do not always return token usage data. When a response carries no usage block, Eneo measures the request it actually sent with LiteLLM’s model-aware token counter (litellm.token_counter), passing the exact message list and tool definitions for that provider call. LiteLLM picks the tokenizer for the model route; when it does not recognise the model it uses its default cl100k_base encoding. If counting fails altogether, Eneo uses a last-resort length estimate (roughly one token per four characters plus a fixed per-message overhead).
The estimate includes:
- Message scaffolding (role framing and per-message overhead) as LiteLLM models it
- Tool and function definitions sent with the request
- Images, priced from their pixel dimensions
This estimation has some known limitations:
- It may over- or undercount tokens since different models use different tokenizers
- Reasoning tokens that do not appear in the response text are not captured
As a result, token counts for self-hosted models should be treated as approximations rather than exact values. When a provider does report a prompt token count, Eneo compares it with its own estimate for the same request and logs a warning if the two drift by more than 20 %, so tokenizer or pricing changes surface in the backend logs instead of silently skewing dashboards.
Data Flow
LLM Provider
│
▼
Eneo Backend (via LiteLLM)
│
├── Provider returns usage ──► Actual token counts recorded
│
└── No usage available ──────► Local estimation used
│
▼
Insights / Admin Dashboard