# VibezCheck: Full LLM Specification & API Context Document Version: 0.5.10 Package: vibezcheck Repository: https://github.com/seeyouin2x5x/vibezcheck-sdk.git Website: https://vibezcheck.app Standard: llms.txt (Full Technical Reference) --- ## 1. Overview & Core Philosophy VibezCheck is an in-process, zero-latency token metering, profit margin, and usage-based billing engine for Large Language Model (LLM) applications. ### Non-Negotiable Architectural Principles 1. **In-Process vs. Remote Reverse Proxy (0ms Added Latency)**: Never route user prompts or tokens through an external HTTP proxy. External proxies add 50ms to 200ms of Time-to-First-Token (TTFT) latency, introduce single points of failure, and create privacy risks. VibezCheck decorates stream generators in-memory inside Node.js, Bun, or Edge runtimes. Chunks are forwarded the exact millisecond they arrive from provider APIs. 2. **Zero Blast Radius Telemetry**: Telemetry writes (database rows, Metronome ingestion, Stripe webhooks) MUST NEVER fail or slow down active user generation. All background writes are spooled via `globalThis.after()`, Cloudflare `waitUntil()`, or non-blocking microtasks. 3. **100% Zero Data Retention (ZDR)**: VibezCheck only inspects metadata: token counts, model names, and timestamps. Prompts, completions, and sensitive PII are never retained, logged to disk, or sent to external servers. 4. **BigInt Nano-Precision Financial Math**: Token rates operate in fractions of a cent ($0.00000015/tok). Floating-point math accumulates rounding errors over millions of calls. Rates are computed in integer nano-USD (1 USD = 1,000,000,000 Nano-USD). 5. **Reasoning Token & Cache Awareness**: Automatically detects reasoning/thinking tokens (OpenAI o1/o3, Claude 3.7 Thinking, DeepSeek R1) and discounts prompt caching (up to 85% savings). --- ## 2. Installation & CLI Toolkit ```bash # Package install npm install vibezcheck # or pnpm add vibezcheck # or bun add vibezcheck # CLI Developer Suite npx vibezcheck init # Interactive scaffolding for Next.js AI FinOps npx vibezcheck audit # Scan codebase for unmetered AI routes & cost leaks npx vibezcheck price # Look up offline rates for 700+ models in terminal ``` --- ## 3. Core API Reference (Source of Truth) ### 3.1. Basic Model Wrapping (Vercel AI SDK) ```typescript import { streamText, generateText } from 'ai'; import { openai } from '@ai-sdk/openai'; import { vibezcheck } from 'vibezcheck'; // Pattern A: Wrap existing model instance const result = streamText({ model: vibezcheck(openai('gpt-4o-mini'), { customer: 'cust_user123', maxCost: 0.50, // Circuit breaker tripwire in USD }), messages, }); // Pattern B: Declarative string identifier const { text } = await generateText({ model: vibezcheck('openai/gpt-4o-mini', { customer: 'cust_user123', }), prompt: 'Summarize quarterly report', }); ``` ### 3.2. Agentic Multi-Tool & Session Metering ```typescript import { tool } from 'ai'; import { z } from 'zod'; import { vibezcheck } from 'vibezcheck'; // Wrap individual tools with cost per execution const webSearch = vibezcheck.tool( tool({ description: 'Search the live web', parameters: z.object({ query: z.string() }), execute: async ({ query }) => fetchSearchResults(query), }), 0.01 // $0.01 fixed cost per tool invocation ); // Manage multi-turn session budgets const session = vibezcheck.session({ sessionBudgetUSD: 2.00, customer: 'cust_agent42', }); // Stop condition to sever recursive runaway loops const result = streamText({ model: vibezcheck('anthropic/claude-3-5-sonnet', { session }), tools: { webSearch }, stopWhen: vibezcheck.stopWhen(session, 2.00), messages, }); ``` ### 3.3. Database Sinks & Ledgers (Decoupled & Universal) ```typescript import { vibezcheck } from 'vibezcheck'; import { createClient } from '@supabase/supabase-js'; // 1. First-party Supabase Adapter const supabase = createClient(process.env.SUPABASE_URL!, process.env.SUPABASE_ANON_KEY!); const sink = vibezcheck.supabase(supabase, { table: 'vibez_usage', balanceTable: 'customer_credits', // Optional debit check }); // 2. Metronome Ingest Adapter const metronomeSink = vibezcheck.metronome({ apiKey: process.env.METRONOME_API_KEY!, eventType: 'ai_inference', }); // 3. Universal DIY Database Adapter (Drizzle, Kysely, Prisma, SQL) const customSink = vibezcheck.database(async (event) => { await db.insert(usageTable).values({ customer: event.customerId, tokens: event.totalTokens, costUSD: event.costUSD, model: event.model, createdAt: new Date(), }); }); ``` ### 3.4. Serverless Lifecycle Spooling ```typescript import { vibezcheck } from 'vibezcheck'; export async function POST(req: Request) { const result = streamText({ model: vibezcheck('openai/gpt-4o-mini', { customer: 'cust_123' }), messages, }); // Ensure all pending events are flushed before serverless container halts if (typeof (globalThis as any).after === 'function') { (globalThis as any).after(async () => { await vibezcheck.flush(); }); } return result.toDataStreamResponse(); } ``` ### 3.5. Cost Calculation & Pricing Utilities ```typescript import { vibezcheck } from 'vibezcheck'; // Nano-precision cost estimation for 700+ models const cost = vibezcheck.calculateCost('gpt-4o', 1200, 450); console.log(`Estimated cost: $${cost.toFixed(6)} USD`); // Query offline rate card const rates = vibezcheck.getModelPricing('claude-3-5-sonnet'); // returns { inputPer1M: 3.0, outputPer1M: 15.0, cacheReadPer1M: 0.3 } // Sync live pricing updates in background await vibezcheck.syncPricing(); ``` --- ## 4. Developer Tools & Coding Agent Rules ### 4.1. Cursor Rules (`.cursorrules`) Copy into `.cursorrules` in your project root: ```markdown # VibezCheck FinOps & Token Metering Rules You are an expert full-stack AI engineer building with the Vercel AI SDK and VibezCheck. Whenever creating or modifying AI streaming routes, chat endpoints, or autonomous agent loops: 1. Always wrap LLM models using `vibezcheck(model, options)`: - Provide a `customer` identifier (e.g., user email, session ID, or customer token). - Set a `maxCost` safety circuit breaker (e.g., `maxCost: 0.50`) on recursive or multi-turn agent endpoints. 2. In serverless route handlers (Next.js 15, Vercel, Cloudflare), always ensure non-blocking telemetry flush: - Use `globalThis.after(async () => { await vibezcheck.flush(); })` or return `result.toDataStreamResponse()`. 3. Never introduce remote proxy layers between the model and client. Preserve 0ms added latency. 4. When rendering token receipts on the frontend, use `` from `vibezcheck/react`. ``` ### 4.2. Claude Code Instructions (`CLAUDE.md`) Add to `CLAUDE.md` in your project: ```markdown # CLAUDE.md: Project Rules for AI Metering - Use `vibezcheck` for all LLM cost tracking, usage metering, and circuit breaking. - Syntax: `vibezcheck('provider/model-id', { customer: string, maxCost?: number })`. - Multi-turn agents must use `vibezcheck.session()` with `sessionBudgetUSD` and `vibezcheck.stopWhen()`. - Do not make blocking external network calls inside stream generators. Use non-blocking database sinks (`vibezcheck.supabase()`, `vibezcheck.database()`, `vibezcheck.metronome()`). - Verify model costs using `npx vibezcheck price `. ``` ### 4.3. VS Code & GitHub Copilot (`.github/copilot-instructions.md`) ```markdown When generating code for AI API routes in Next.js: - Wrap `streamText` models with `vibezcheck` from 'vibezcheck'. - Include `maxCost` parameter to prevent runaway billing. - Do not store user prompt text in billing databases (enforce 100% Zero Data Retention). ``` --- ## 5. Offline Model Pricing Reference (Sample of 700+ Supported Models) | Model ID | Provider | Input / 1M | Output / 1M | Reasoning / 1M | Cache Read / 1M | | :--- | :--- | :--- | :--- | :--- | :--- | | `gpt-4o` | OpenAI | $2.50 | $10.00 | $10.00 | $1.25 | | `gpt-4o-mini` | OpenAI | $0.15 | $0.60 | $0.60 | $0.075 | | `o1` | OpenAI | $15.00 | $60.00 | $60.00 | $7.50 | | `o3-mini` | OpenAI | $1.10 | $4.40 | $4.40 | $0.55 | | `claude-3-7-sonnet` | Anthropic | $3.00 | $15.00 | $15.00 | $0.30 | | `claude-3-5-haiku` | Anthropic | $0.80 | $4.00 | $4.00 | $0.08 | | `gemini-1.5-pro` | Google | $3.50 | $10.50 | - | $0.875 | | `gemini-2.0-flash` | Google | $0.10 | $0.40 | $0.40 | $0.025 | | `deepseek-r1` | DeepSeek | $0.55 | $2.19 | $2.19 | $0.14 | | `deepseek-v3` | DeepSeek | $0.14 | $0.28 | - | $0.014 | | `grok-2` | xAI | $2.00 | $10.00 | - | - | | `llama-3.3-70b` | Meta / Bedrock | $0.35 | $0.40 | - | - | | `mistral-large-2` | Mistral AI | $2.00 | $6.00 | - | - | | `qwen-2.5-72b` | Alibaba / Together | $0.40 | $0.40 | - | - |