Documentation

DoCoreAI Developer Docs

Everything you need to install, configure, and run DoCoreAI in production.

New to LLM monitoring? Start with what is LLM observability or see our LLM monitoring tools comparison.


v2.1.0
Python 3.12+
Phase 2.2 · Jun 2026
$ pip install docoreai
Copied!

Quick Start

DoCoreAI installs in the same Python environment as your application and starts intercepting all LLM SDK calls automatically. No changes to your existing code required.

ℹ️
Requirements: Python 3.12+ · pip · A free org token from docoreai.com . Actively tested on Windows. macOS and Linux should work — please report any issues.
1

Install DoCoreAI

Install in the same environment as your application. If pip is not in PATH (e.g. Windows + Python 3.13), use the python -m prefix.

bash
# Standard
pip install docoreai

# If pip not in PATH (Windows/Python 3.13)
python -m pip install docoreai
2

Configure with your org token

Generate your free org token at docoreai.com then run the one-time setup. Your token starts with to_.

bash
docoreai config
💡
How to get your token: Individual tokens are generated at docoreai.com/generate-token . For organisation-level setup and multi-account management, visit docoreai.com/account-settings . Paste your token into the Telemetry Token field in the configuration window and click Save. All settings are stored locally on your machine.
3

Start DoCoreAI

Run alongside your application. DoCoreAI automatically intercepts all LLM SDK calls in your environment. Stop with Ctrl+C — active requests complete before exit.

bash
# Start
docoreai start

# Stop (graceful shutdown)
Ctrl+C
✅
That's it. No code changes needed. DoCoreAI monkey-patches all active LLM SDKs automatically at startup. Your application continues to work exactly as before.
4

Dev Testing — Multiple Clients in venv

When testing multiple client environments locally (e.g. HSBC and Target), give each client its own virtual environment. This isolates the .pth file, SQLite DB, and LightGBM model per client.

bash
# Create a venv per client
python -m venv hsbc\venv
python -m venv target\venv

# Install DoCoreAI into each venv
hsbc\venv\Scripts\pip install docoreai
target\venv\Scripts\pip install docoreai

# Start each client — cd into the client folder first
cd hsbc
..\hsbc\venv\Scripts\docoreai start

# In a second terminal
cd target
..\target\venv\Scripts\docoreai start

# Run client apps using their venv's Python directly
hsbc\venv\Scripts\python.exe hsbc\hsbc.py
target\venv\Scripts\python.exe target\target.py
ℹ️
Why separate venvs? Each venv has its own site-packages folder, so the docoreai_autopatch.pth file is fully isolated per client. Two clients can run simultaneously without any path or DB collision.
5

Prod Deployment — Container Setup

In production, DoCoreAI is installed inside the client app's container alongside the application. One environment variable tells DoCoreAI where the client environment lives. No cd ritual, no hardcoded paths.

dockerfile
# Dockerfile
FROM python:3.12-slim
WORKDIR /app

# Required — tells DoCoreAI where the environment lives
ENV DOCOREAI_ENV_PATH=/app

COPY requirements.txt .
RUN pip install -r requirements.txt
RUN pip install docoreai

COPY . .

# Start DoCoreAI then run your app
CMD ["sh", "-c", "docoreai start & python your_app.py"]
📌
DOCOREAI_ENV_PATH is the only required deployment variable. Set it to the directory where your app runs. DoCoreAI uses this to locate the SQLite database and resolve all runtime paths — no .env file needed in production.

How It Works

DoCoreAI runs as a sidecar in the same Python environment as your application. It never sits in your call path. Your app calls the LLM directly and unchanged. DoCoreAI listens via SDK patch, extracts metadata locally, and sends only cost and token data to the cloud.

Request Flow

On every LLM call, DoCoreAI performs these steps autonomously — before and after the actual API call:

  • Governance check — risky prompt detection, PII scan at the edge
  • Token prediction — ML model predicts how many tokens the response actually needs
  • Budget check — validates request against daily spend limit
  • Pacing adjustment — compares actual vs. expected spend rate, throttles if over pace
  • Soft limit injection — adds concise guidance to the prompt when budget is tight
  • LLM call executed — your app calls the provider directly, unchanged
  • Post-processing — compares prediction vs. actual, logs metadata, updates learning models

Privacy Model

🔒
Prompts never leave your network. DoCoreAI extracts only cost and token metadata locally. No prompt content, no response content, no PII is ever sent to the cloud. Local telemetry is stored in SQLite on your machine.
Data Type Stored Locally Sent to Cloud
Prompt content ✕ Never ✕ Never
Response content ✕ Never ✕ Never
Token counts ✓ Yes ✓ Aggregated
Cost per request ✓ Yes ✓ Aggregated
Latency ✓ Yes ✓ Aggregated
Model name ✓ Yes ✓ Aggregated

Supported Traffic Shapes

Which GenAI architecture patterns DoCoreAI can currently instrument, and what is on the roadmap.

Architecture Supported Stage Target
Single-shot ✓ Yes Testing 100% complete
Agentic tool-calling loop ✓ Yes Development Sept 2026
Multi-turn conversational ― TBD — —
Multi-agent orchestration ✓ Yes Research Feb 2027
Large-context single-shot (RAG) ✓ Yes Not Started July 2027

Roadmap targets are indicative and subject to change. Multi-turn conversational support is under evaluation — behaviour with stateful session patterns is not yet defined.

Budget Control

Set a daily spend limit and choose how DoCoreAI enforces it. Six enforcement modes available — from intelligent optimization to hard block.

Budget Modes

Mode Behaviour Recommended
smart_reduce Intelligently reduces token limits as budget tightens. Service never stops. ✓ Default
reduce_tokens Linear token reduction based on remaining budget percentage. Most cases
warn_and_allow Logs warnings but never blocks. Good for learning phase. Week 1
block Hard stop when budget is exhausted. Strict compliance use cases. Compliance
use_fallback Switches to cheaper model when budget is tight. Multi-model setups. Advanced
ignore No enforcement. Development only. Dev only
python
# Budget configuration example
BUDGET_CONFIG = {
    'daily_budget': 100.00,
    'mode': 'smart_reduce',
    'use_soft_limits': True,
    'use_pacing_engine': True,
    'allow_user_override': False
}

Pacing Engine

The pacing engine spreads your daily budget evenly across 24 hours — like Google Ads pacing. It detects when you are spending faster than expected and applies graduated throttling automatically. No service interruptions. No manual intervention.

Pacing Strategies

Strategy How It Works Use When
even Divides budget equally across 24 hours. Simple linear model. First 7 days
adaptive Learns your hourly usage patterns over 7 days, then paces to your real profile. ✓ Recommended after learning
peak_aware Hybrid — adaptive baseline plus real-time spike detection. Allows temporary over-pace during detected events. Event-driven traffic
💡
Learning period: DoCoreAI starts with the even strategy and automatically switches to adaptive after 7 days and 1,000+ requests. No action needed.

Soft Limits

Instead of hard-truncating responses, DoCoreAI injects a concise guidance message into the system prompt — gently guiding the LLM to shorter responses when budget is tight. Quality is preserved. Responses are never cut off mid-sentence.

When Soft Limits Activate

  • Budget remaining is below 20%
  • Single request cost exceeds $1.00
  • Request cost exceeds remaining budget
  • always mode is enabled in config
python
SOFT_LIMIT_CONFIG = {
    'enabled': True,
    'buffer_pct': 20,
    'min_words': 50,
    'hard_cap_multiplier': 1.2,
    'use_when': {
        'budget_tight': True,
        'high_cost_request': True
    }
}

Governance & PII Detection

DoCoreAI performs governance checks on every request before the LLM call is made. PII detection runs at the edge — inside your environment, before any data leaves your network.

  • PII detection — scans for names, emails, phone numbers, credit cards before API call
  • Risky prompt detection — flags or blocks prompts matching governance rules
  • Block vs. warn mode — configurable per governance rule
  • Audit trail — all governance events logged locally, no sensitive content ever sent to cloud
⚠️
SOC2 certification in progress. Governance features are production-ready. Formal certification is targeted for Q3 2026.

Prediction Engine

The prediction engine uses a LightGBM model to estimate how many tokens each response will actually need — replacing the wasteful default ceiling with a precise prediction.

How Predictions Work

  • Extracts 30+ features from each request (prompt tokens, model, temperature, user history)
  • LightGBM model predicts completion tokens needed
  • Applies 1.05× safety margin
  • Caps at model's max output tokens
  • Default max_tokens of 2,000+ replaced with precise estimate — typically 200–400 tokens
📊
Projected savings: At 30,000 requests/month, prediction alone saves approximately $990/month before pacing or soft limits contribute additional savings.

Auto-Retraining

When prediction accuracy degrades, DoCoreAI automatically detects drift, retrains the LightGBM model on recent telemetry, runs an A/B test against the previous champion, and promotes the better model. No action required from your team.

Retrain Triggers

Trigger Default Value
days_since_last_training 30 days
min_new_predictions 1,000 samples
drift_detected MAE increase > 20%
cooldown_hours 24 hours between retrains
max_retrains_per_week 2 per week

Supported Providers

DoCoreAI detects and wraps all active SDKs automatically at startup. No provider-specific configuration required.

OpenAI
GPT-4, GPT-4 Turbo, GPT-4o, GPT-3.5 Turbo and all variants
Auto-detected
Anthropic
Claude 3 family, Claude 3.5, and newer models
Auto-detected
Google Gemini
Gemini Pro, Gemini Ultra, Gemini Flash
Auto-detected
Groq
Llama, Mixtral, and all Groq-hosted models
Auto-detected
AWS Bedrock
All supported foundation models via Bedrock API
Auto-detected
Ollama
Local model deployments via Ollama runtime
Auto-detected

CLI Reference

All DoCoreAI operations are available via the command line.

Command Description
docoreai start Start DoCoreAI alongside your application
docoreai stop Graceful shutdown — active requests complete first
docoreai config One-time setup with your org token
docoreai train-tokens-model Manually trigger model retraining
docoreai sync-pricing Sync latest LLM pricing from server
docoreai start --env=dev Start in development environment mode
💡
Windows / Python 3.13: If docoreai is not found in PATH, prefix all commands with python -m docoreai.

Configuration Reference

All configuration lives in phase2_config.py. Key settings grouped by feature area.

python
# phase2_config.py — key settings

# Budget
BUDGET_CONFIG = {
    'daily_budget': 100.00,
    'mode': 'smart_reduce',
    'use_soft_limits': True,
    'use_pacing_engine': True
}

# Pacing
PACING_ENGINE_CONFIG = {
    'enabled': True,
    'strategy': 'adaptive',
    'throttle_threshold_pct': 20,
    'aggressive_threshold_pct': 40
}

# Auto-Retrain
AUTO_RETRAIN_CONFIG = {
    'auto_trigger': True,
    'cooldown_hours': 24,
    'max_retrains_per_week': 2,
    'min_mae_increase_pct': 20.0
}

# A/B Testing
AB_TESTING_CONFIG = {
    'min_test_duration_hours': 24,
    'challenger_traffic_pct': 0.10,
    'auto_promote': True,
    'auto_rollback': True
}

Troubleshooting

Common issues and their solutions. Click any item to expand.

docoreai start fails or command not found +
Check Python version: Requires Python 3.12+. Run python --version to verify.

pip not in PATH (Windows/3.13): Use python -m pip install docoreai and python -m docoreai start.

Missing dependencies: Run pip install -r requirements.txt.
Predictions returning None +
No trained model found in saved_models/. Run docoreai train-tokens-model to train the initial model. Requires 5,000+ telemetry samples. During the learning phase, safe estimates are used automatically.
Budget shows $0.00 or incorrect amount +
Check autopatch is active: Verify DoCoreAI started correctly and is intercepting calls.

Timezone mismatch: Budget resets at server midnight (UTC). If your timezone differs, the reset will appear at a different local time.

Missing pricing: Run docoreai sync-pricing to update the local model registry.
Pacing not activating despite being over budget pace +
Check that throttle_threshold_pct is not set too high in phase2_config.py. Default is 20% — if set to 80%+, pacing will not trigger until severely over pace.

Also verify enabled: True in PACING_ENGINE_CONFIG.
A/B test running for more than 7 days with no decision +
Insufficient data — each model needs 100+ predictions for comparison. At low volume, tests may take longer to reach the minimum sample threshold.

To manually promote or rollback:
from docore_ai.train.ab_test_controller import promote_challenger
promote_challenger( test_id='ab_YYYYMMDD_HHMMSS')

FAQ

Does DoCoreAI sit in my API call path? +
No. DoCoreAI monkey-patches the LLM SDK locally and listens before and after each call in the same Python process. Your app calls the LLM provider directly, unchanged. DoCoreAI never proxies, routes, or intercepts the actual network call.
Are my prompts stored anywhere? +
Never. Prompt content and response content are never stored locally or sent to the cloud. Only cost, token counts, latency, and model name are collected — all aggregated before leaving your environment.
Do I need to change my existing code? +
No. DoCoreAI auto-patches all active LLM SDK calls at startup. Three commands to integrate: pip install docoreai, docoreai config, docoreai start. Your existing code is untouched.
What happens if DoCoreAI crashes? +
Your application continues to work normally. DoCoreAI is a sidecar process — if it stops, your LLM calls continue unaffected. You simply lose the cost tracking and budget control features until it is restarted.
How long does the learning phase take? +
Budget tracking and governance work immediately from day one. The prediction model requires 5,000+ telemetry samples and 7+ days before switching from even to adaptive pacing. At 200 requests/day this is approximately 25 days. At 1,000/day, approximately 5 days.
Is v2.1.0 production-ready? +
Yes. v2.1.0 is the current stable release, built on 17 months of research and validated against waste patterns across 20+ enterprise AI deployments. The platform is prototype stage in terms of cloud dashboard features, but the core SDK — budget control, pacing, governance, and prediction — is production-ready. Real usage data helps improve prediction models faster. Enterprise pilots and design partners are actively welcomed.

Need help or want to go deeper?

Join as a design partner for white-glove founder support and direct configuration help. Or reach Saji directly.

-->
Scroll to Top