DoCoreAI Developer Docs
Everything you need to install, configure, and run DoCoreAI in production.
New to LLM monitoring? Start with what is LLM observability or see our LLM monitoring tools comparison.
$ pip install docoreai
Quick Start
DoCoreAI installs in the same Python environment as your application and starts intercepting all LLM SDK calls automatically. No changes to your existing code required.
Install DoCoreAI
Install in the same environment as your
application. If pip is not in PATH
(e.g. Windows + Python 3.13),
use the python -m prefix.
# Standard
pip install docoreai
# If pip not in PATH (Windows/Python 3.13)
python -m pip install docoreai
Configure with your org token
Generate your free org token at
docoreai.com then run
the one-time setup. Your token starts
with to_.
docoreai config
Start DoCoreAI
Run alongside your application. DoCoreAI
automatically intercepts all LLM SDK calls
in your environment. Stop with
Ctrl+C — active requests
complete before exit.
# Start
docoreai start
# Stop (graceful shutdown)
Ctrl+C
Dev Testing — Multiple Clients in venv
When testing multiple client environments
locally (e.g. HSBC and Target), give each
client its own virtual environment. This
isolates the .pth file,
SQLite DB, and LightGBM model per client.
# Create a venv per client
python -m venv hsbc\venv
python -m venv target\venv
# Install DoCoreAI into each venv
hsbc\venv\Scripts\pip install docoreai
target\venv\Scripts\pip install docoreai
# Start each client — cd into the client folder first
cd hsbc
..\hsbc\venv\Scripts\docoreai start
# In a second terminal
cd target
..\target\venv\Scripts\docoreai start
# Run client apps using their venv's Python directly
hsbc\venv\Scripts\python.exe hsbc\hsbc.py
target\venv\Scripts\python.exe target\target.py
site-packages folder,
so the docoreai_autopatch.pth
file is fully isolated per client.
Two clients can run simultaneously
without any path or DB collision.
Prod Deployment — Container Setup
In production, DoCoreAI is installed inside
the client app's container alongside the
application. One environment variable tells
DoCoreAI where the client environment lives.
No cd ritual, no hardcoded paths.
# Dockerfile
FROM python:3.12-slim
WORKDIR /app
# Required — tells DoCoreAI where the environment lives
ENV DOCOREAI_ENV_PATH=/app
COPY requirements.txt .
RUN pip install -r requirements.txt
RUN pip install docoreai
COPY . .
# Start DoCoreAI then run your app
CMD ["sh", "-c", "docoreai start & python your_app.py"]
DOCOREAI_ENV_PATH is the
only required deployment variable.
Set it to the directory where your app
runs. DoCoreAI uses this to locate the
SQLite database and resolve all runtime
paths — no .env file needed
in production.
How It Works
DoCoreAI runs as a sidecar in the same Python environment as your application. It never sits in your call path. Your app calls the LLM directly and unchanged. DoCoreAI listens via SDK patch, extracts metadata locally, and sends only cost and token data to the cloud.
Request Flow
On every LLM call, DoCoreAI performs these steps autonomously — before and after the actual API call:
- Governance check — risky prompt detection, PII scan at the edge
- Token prediction — ML model predicts how many tokens the response actually needs
- Budget check — validates request against daily spend limit
- Pacing adjustment — compares actual vs. expected spend rate, throttles if over pace
- Soft limit injection — adds concise guidance to the prompt when budget is tight
- LLM call executed — your app calls the provider directly, unchanged
- Post-processing — compares prediction vs. actual, logs metadata, updates learning models
Privacy Model
| Data Type | Stored Locally | Sent to Cloud |
|---|---|---|
| Prompt content | ✕ Never | ✕ Never |
| Response content | ✕ Never | ✕ Never |
| Token counts | ✓ Yes | ✓ Aggregated |
| Cost per request | ✓ Yes | ✓ Aggregated |
| Latency | ✓ Yes | ✓ Aggregated |
| Model name | ✓ Yes | ✓ Aggregated |
Supported Traffic Shapes
Which GenAI architecture patterns DoCoreAI can currently instrument, and what is on the roadmap.
| Architecture | Supported | Stage | Target |
|---|---|---|---|
Single-shot |
✓ Yes | Testing | 100% complete |
Agentic tool-calling loop |
✓ Yes | Development | Sept 2026 |
Multi-turn conversational |
― TBD | — | — |
Multi-agent orchestration |
✓ Yes | Research | Feb 2027 |
Large-context single-shot (RAG) |
✓ Yes | Not Started | July 2027 |
Roadmap targets are indicative and subject to change. Multi-turn conversational support is under evaluation — behaviour with stateful session patterns is not yet defined.
Budget Control
Set a daily spend limit and choose how DoCoreAI enforces it. Six enforcement modes available — from intelligent optimization to hard block.
Budget Modes
| Mode | Behaviour | Recommended |
|---|---|---|
| smart_reduce | Intelligently reduces token limits as budget tightens. Service never stops. | ✓ Default |
| reduce_tokens | Linear token reduction based on remaining budget percentage. | Most cases |
| warn_and_allow | Logs warnings but never blocks. Good for learning phase. | Week 1 |
| block | Hard stop when budget is exhausted. Strict compliance use cases. | Compliance |
| use_fallback | Switches to cheaper model when budget is tight. Multi-model setups. | Advanced |
| ignore | No enforcement. Development only. | Dev only |
# Budget configuration example
BUDGET_CONFIG = {
'daily_budget': 100.00,
'mode': 'smart_reduce',
'use_soft_limits': True,
'use_pacing_engine': True,
'allow_user_override': False
}
Pacing Engine
The pacing engine spreads your daily budget evenly across 24 hours — like Google Ads pacing. It detects when you are spending faster than expected and applies graduated throttling automatically. No service interruptions. No manual intervention.
Pacing Strategies
| Strategy | How It Works | Use When |
|---|---|---|
| even | Divides budget equally across 24 hours. Simple linear model. | First 7 days |
| adaptive | Learns your hourly usage patterns over 7 days, then paces to your real profile. | ✓ Recommended after learning |
| peak_aware | Hybrid — adaptive baseline plus real-time spike detection. Allows temporary over-pace during detected events. | Event-driven traffic |
even strategy and automatically
switches to adaptive after
7 days and 1,000+ requests. No action needed.
Soft Limits
Instead of hard-truncating responses, DoCoreAI injects a concise guidance message into the system prompt — gently guiding the LLM to shorter responses when budget is tight. Quality is preserved. Responses are never cut off mid-sentence.
When Soft Limits Activate
- Budget remaining is below 20%
-
Single request cost exceeds
$1.00 - Request cost exceeds remaining budget
-
alwaysmode is enabled in config
SOFT_LIMIT_CONFIG = {
'enabled': True,
'buffer_pct': 20,
'min_words': 50,
'hard_cap_multiplier': 1.2,
'use_when': {
'budget_tight': True,
'high_cost_request': True
}
}
Governance & PII Detection
DoCoreAI performs governance checks on every request before the LLM call is made. PII detection runs at the edge — inside your environment, before any data leaves your network.
- PII detection — scans for names, emails, phone numbers, credit cards before API call
- Risky prompt detection — flags or blocks prompts matching governance rules
- Block vs. warn mode — configurable per governance rule
- Audit trail — all governance events logged locally, no sensitive content ever sent to cloud
Prediction Engine
The prediction engine uses a LightGBM model to estimate how many tokens each response will actually need — replacing the wasteful default ceiling with a precise prediction.
How Predictions Work
- Extracts 30+ features from each request (prompt tokens, model, temperature, user history)
- LightGBM model predicts completion tokens needed
- Applies 1.05× safety margin
- Caps at model's max output tokens
- Default max_tokens of 2,000+ replaced with precise estimate — typically 200–400 tokens
Auto-Retraining
When prediction accuracy degrades, DoCoreAI automatically detects drift, retrains the LightGBM model on recent telemetry, runs an A/B test against the previous champion, and promotes the better model. No action required from your team.
Retrain Triggers
| Trigger | Default Value |
|---|---|
| days_since_last_training | 30 days |
| min_new_predictions | 1,000 samples |
| drift_detected | MAE increase > 20% |
| cooldown_hours | 24 hours between retrains |
| max_retrains_per_week | 2 per week |
Supported Providers
DoCoreAI detects and wraps all active SDKs automatically at startup. No provider-specific configuration required.
CLI Reference
All DoCoreAI operations are available via the command line.
| Command | Description |
|---|---|
| docoreai start | Start DoCoreAI alongside your application |
| docoreai stop | Graceful shutdown — active requests complete first |
| docoreai config | One-time setup with your org token |
| docoreai train-tokens-model | Manually trigger model retraining |
| docoreai sync-pricing | Sync latest LLM pricing from server |
| docoreai start --env=dev | Start in development environment mode |
docoreai is not found in
PATH, prefix all commands with
python -m docoreai.
Configuration Reference
All configuration lives in
phase2_config.py.
Key settings grouped by feature area.
# phase2_config.py — key settings
# Budget
BUDGET_CONFIG = {
'daily_budget': 100.00,
'mode': 'smart_reduce',
'use_soft_limits': True,
'use_pacing_engine': True
}
# Pacing
PACING_ENGINE_CONFIG = {
'enabled': True,
'strategy': 'adaptive',
'throttle_threshold_pct': 20,
'aggressive_threshold_pct': 40
}
# Auto-Retrain
AUTO_RETRAIN_CONFIG = {
'auto_trigger': True,
'cooldown_hours': 24,
'max_retrains_per_week': 2,
'min_mae_increase_pct': 20.0
}
# A/B Testing
AB_TESTING_CONFIG = {
'min_test_duration_hours': 24,
'challenger_traffic_pct': 0.10,
'auto_promote': True,
'auto_rollback': True
}
Troubleshooting
Common issues and their solutions. Click any item to expand.
docoreai start fails
or command not found
+
python --version to verify.
pip not in PATH (Windows/3.13): Use
python -m pip install docoreai
and python -m docoreai start.
Missing dependencies: Run
pip install -r requirements.txt.
None
+
saved_models/.
Run docoreai train-tokens-model
to train the initial model.
Requires 5,000+ telemetry samples.
During the learning phase, safe estimates
are used automatically.
Timezone mismatch: Budget resets at server midnight (UTC). If your timezone differs, the reset will appear at a different local time.
Missing pricing: Run
docoreai sync-pricing
to update the local model registry.
throttle_threshold_pct
is not set too high in
phase2_config.py.
Default is 20% — if set to 80%+,
pacing will not trigger until severely
over pace.
Also verify
enabled: True
in PACING_ENGINE_CONFIG.
To manually promote or rollback:
from docore_ai.train.ab_test_controller
import promote_challenger
promote_challenger(
test_id='ab_YYYYMMDD_HHMMSS')
FAQ
pip install docoreai,
docoreai config,
docoreai start.
Your existing code is untouched.
even to adaptive
pacing. At 200 requests/day this is
approximately 25 days. At 1,000/day,
approximately 5 days.
Need help or want to go deeper?
Join as a design partner for white-glove founder support and direct configuration help. Or reach Saji directly.
