How to Track LLM API Costs and Token Spend Without Storing Prompts

How to Track LLM API Costs and Token Spend Without Storing Prompts

Most LLM cost monitoring platforms work by intercepting requests and forwarding them through an external cloud-hosted proxy, introducing severe data governance risks. This step-by-step developer tutorial demonstrates how to track your LLM API costs completely within a standard Python environment using local runtime instrumentation and genuine zero prompt storage.

The Compliance Problem with Cloud Gateways

For organizations operating under strict internal security policies or regulatory frameworks like HIPAA, GDPR, and SOC 2, routing production payload wrappers outside the cluster boundary creates major vulnerabilities. Every prompt sent to an external proxy turns into a data leak liability.

Fortunately, tracking your runtime infrastructure expenses does not require capturing raw text. By implementing local metadata extraction inside your application environment, you can evaluate model metrics, performance drops, and billing overhead without sending sensitive strings to a third party.

Strategic Resource: This local-first implementation forms a core architectural component of a broader operational design. To understand the complete model framework, read our comprehensive foundational guide: The Ultimate Guide to Privacy-First LLM Observability.

The Anatomy of Local Metadata Extraction

To accurately monitor token spend, your telemetry pipeline doesn't need text access. It only requires a concise, isolated payload dictionary extracted dynamically at the response lifecycle point:

Telemetry Field Purpose & Governance Insight
input_tokens Tracks core prompt payload sizes and RAG pipeline context growth.
output_tokens Measures completion response lengths and catches script generation loops.
model_name Identifies exactly which provider model (e.g., gpt-4o, claude-3-5-sonnet, o1-pro, gemini-2.5-pro) generated the cost events.
latency_ms Measures end-to-end execution durations to monitor backend degradation.
status_code Logs exact response state codes to analyze system exception rates.

At no point during extraction does the local engine capture structural text markers like prompt, messages, or user_content. It isolates metrics, discarding the text arrays entirely before logging.

Step-by-Step Guide: Implementing Python LLM Tracking

Step 1: Install the Dependency and Configure Environment Variables

Pull the telemetry package down via your native installer and define your environmental workspace variables to specify local processing flags:

pip install docoreai
DOCOREAI_ENABLE=true
DOCOREAI_LOG_ONLY=true
DOCOREAI_API_URL=http://localhost:8000
DOCOREAI_TOKEN=local_runtime_token

Note: DoCoreAI version 2+ does not allow direct usage of .env file. The configurations are done with easy UI screens. Use command `docoreai config`

Step 2: Initialize Runtime Monitoring

Traditional tracking requires boilerplate logging commands wrapped around every API call. With runtime monkey-patching, you trigger initialization once at application startup, letting the module securely capture response states in the background:

import docoreai
from openai import OpenAI

# Initialize the telemetry hook before any SDK calls
docoreai.start()
client = OpenAI()

response = client.chat.completions.create(
    model="gpt-4o",
    messages=[{"role": "user", "content": "Explain vector embeddings."}]
)
print(response.choices[0].message.content)

Step 3: Simulate Multi-Turn Workloads

Run sequential queries to confirm how metrics build up across independent loops:

questions = ["What is RAG?", "Explain vector search.", "How does semantic retrieval work?"]

for question in questions:
    response = client.chat.completions.create(
        model="gpt-4o",
        messages=[{"role": "user", "content": question}]
    )
    print(response.choices[0].message.content[:100])

Verifying Data Governance in Your Local Database

Security assertions should always be audited. You can verify that your pipeline maintains absolute zero prompt storage by inspecting the local SQLite telemetry file inside your application root path.

Open the workspace database via your command terminal interface:

sqlite3 telemetry.db

Pull up the event logs schema directly:

.schema telemetry_events

The system returns the explicit column schema definition:

CREATE TABLE telemetry_events (
    id INTEGER PRIMARY KEY,
    input_tokens INTEGER,
    output_tokens INTEGER,
    total_tokens INTEGER,
    model_name TEXT,
    latency_ms INTEGER,
    status_code INTEGER,
    created_at TEXT
);

Notice the complete structural exclusion of columns like prompt, messages, or response_text. Only scalar numerical parameters, timestamps, and model string flags exist in the storage layout.

Maintain Visibility Without Compliance Risk

Balancing financial observability against rigorous enterprise data protection guidelines doesn't require proxy middleman layers. By tracking metadata at the runtime client boundary, you protect budgets without exposing corporate data assets.

Ready to Monitor LLM Infrastructure Safely?

Deploy secure, local-first token tracking and autonomous cost governance into your application in minutes.

-->
Scroll to Top