AI usage metering for applications.
Token Metering vs. Dollar Metering
Raw tokens cannot be added together across different models. 100k tokens of GPT-4o-mini is $0.015, while 100k tokens of Claude 3.5 Sonnet is $0.30, and 100k reasoning tokens of o1 is $6.00. Token counts alone leave finance and product teams blind.
VibezCheck calculates exact micro-dollar provider cost and customer retail price at runtime. Usage events are immediately compatible with invoicing engines, database ledgers, customer balances, and margin analytics.
Multi-Dimensional Telemetry Dimensions
Captures prompt, completion, reasoning, and cached tokens per invocation with sub-millisecond precision.
Binds requests to user IDs, organizations, enterprise workspaces, and subscription tier plans.
Tag telemetry by feature (e.g. smart-search, code-generation, auto-summary) to track feature-level unit economics.
Application-Layer Metering in 1 Line
Zero Proxy Overheadimport { streamText } from 'ai';
import { vibezcheck } from 'vibezcheck';
// Wrap any Vercel AI SDK model with rich customer & feature metadata
const result = streamText({
model: vibezcheck('anthropic/claude-3-5-sonnet', {
customer: {
id: 'org_acme_corp',
plan: 'enterprise_annual',
},
metadata: {
feature: 'document_intelligence',
environment: 'production',
},
onUsage: async (event) => {
// Non-blocking telemetry event emitted as soon as stream concludes
console.log('Provider Cost: $' + event.cost.totalUSD);
console.log('Prompt Tokens: ' + event.usage.promptTokens);
console.log('Cached Tokens: ' + event.usage.cachedPromptTokens);
// Persist to your database or billing platform
await db.aiUsageEvents.insert({
customerId: event.customerId,
model: event.model,
costUSD: event.cost.totalUSD,
tokens: event.usage.totalTokens,
timestamp: new Date(event.timestamp),
});
},
}),
prompt: 'Analyze this balance sheet and identify key capital expenditures',
});Why in-process metering wins for production
- 0ms added network latency
- No third-party proxy point of failure
- Provider credentials stay safely in your environment
- Runs natively in Node, Bun, and Edge functions
- Adds 30ms – 150ms network hop on every request
- Proxy downtime takes down your entire AI application
- Requires handing master API keys to an external vendor
- Regional routing and egress bandwidth costs
Frequently Asked Questions
What is AI usage metering?
AI usage metering is the automated measurement of computational resources consumed by AI requests—including prompt tokens, completion tokens, prompt cache discounts, and external tool calls—denominated in both token units and real dollar costs.
How does in-process metering compare to an AI proxy or gateway?
An AI proxy sits between your server and the LLM provider, adding a network hop (often 30ms to 150ms+ of latency), creating a single point of operational failure, and requiring you to share provider API keys. VibezCheck operates in-process inside your TypeScript runtime, inspecting native stream chunks with 0ms network latency and zero credential sharing.
How does VibezCheck handle prompt caching and reasoning tokens?
VibezCheck reads raw provider usage objects (including OpenAI cached_tokens/completion_tokens_details and Anthropic cache_read_input_tokens). It applies accurate cache discounts (typically 50% to 90% cheaper) and rates reasoning tokens according to provider tier cards.
What database adapters are supported out of the box?
VibezCheck provides built-in adapters for Supabase (createSupabaseAdapter), Metronome (createMetronomeAdapter), and custom Postgres/SQL storage via createDatabaseAdapter. You can also listen to events directly via the onUsage callback.
Meter your application usage today
Zero network latency. Sub-cent accuracy. Instant customer attribution.