Skip to main content

Mounting the Python Backend and Scaling It

The tokentrap package is a normal FastAPI application factory. That means it mounts into an app you already run, picks up configuration from environment variables, and - when one process stops being enough - hands you a two-method interface to put sessions in Redis.

Mount into your existing app​

from fastapi import FastAPI
from token_trap import create_app, TrapConfig

app = FastAPI()

trap = create_app(TrapConfig(
persona="Acme Support AI",
strength="maximum",
canary_tokens=("acme-audit-3131",),
rate_limit_per_minute=20,
))
app.mount("/trap", trap)
# POST /trap/api/chat -> widget protocol
# POST /trap/v1/chat/completions -> OpenAI bait
# GET /trap/api/healthz -> liveness

Everything lands under your prefix with CORS already open, so a widget anywhere on your domain can target it:

TokenTrap.init({ container: "#trap", apiEndpoint: "https://api.example.com/trap" });

Prefer to keep the honeypot far from your real app? Run it standalone - same process model either way:

tokentrap serve --host 0.0.0.0 --port 8787 --strength aggressive --canary-tokens prod-canary-1

Configuration from the environment​

Every TrapConfig field has a TOKENTRAP_* counterpart via TrapConfig.from_env_prefix(None), which is what the CLI reads before applying flags:

Env varFeeds fieldNotes
TOKENTRAP_PERSONApersonastring
TOKENTRAP_STRENGTHstrengthvalidated against moderate/aggressive/maximum
TOKENTRAP_CANARY_TOKENScanary_tokenscomma-separated
TOKENTRAP_EXTRA_KEYWORDSextra_keywordscomma-separated
TOKENTRAP_OVERRIDE_KEYWORDSoverride_keywords1 / true / yes
TOKENTRAP_RATE_LIMITrate_limit_per_minuteint
TOKENTRAP_SESSION_TTLsession_ttl_secondsint
TOKENTRAP_LOG_WEBHOOKlog_webhookURL
TOKENTRAP_LLM_MODELllm_modelrequires the llm extra

That makes container deployments boring: no config file, just env. CLI flags win over env vars where both are present.

What's per-process (and what to do about it)​

Two pieces of state live in process memory by default:

  1. Sessions (InMemorySessionStore) - thread-safe dict with TTL expiry (24h default) and opportunistic sweeping past 10,000 records. Tracks turn counts and last meta only; transcripts are never stored server-side.
  2. Rate limiting (SlidingWindowRateLimiter) - 60-second sliding windows keyed by client IP (leftmost x-forwarded-for entry preferred), stale-key pruning past 5,000 tracked IPs.

Run uvicorn main:app --workers 4 behind a load balancer and each worker keeps its own session map. Consequences: escalation depth can reset when requests bounce between workers, and rate limits are enforced per worker rather than globally.

Swapping in Redis-backed sessions​

Both backends share the same tiny session contract. In Python it's three methods on the store:

  • get_or_create(session_id) -> record with .turns
  • increment_turns(session_id) -> new count
  • set_last_meta(session_id, meta_dict) -> None

A minimal Redis implementation of that interface:

import json
import redis
from token_trap.session import SessionRecord

class RedisSessionStore:
def __init__(self, url: str, ttl_seconds: int = 86_400) -> None:
self._r = redis.Redis.from_url(url)
self._ttl = ttl_seconds

def _key(self, session_id: str) -> str:
return f"tokentrap:session:{session_id}"

def get_or_create(self, session_id: str) -> SessionRecord:
raw = self._r.get(self._key(session_id))
if raw is not None:
data = json.loads(raw)
rec = SessionRecord(session_id)
rec.turns = data["turns"]
return rec
fresh = SessionRecord(session_id)
self._r.set(self._key(session_id), json.dumps({"turns": 0}), ex=self._ttl)
return fresh

def increment_turns(self, session_id: str) -> int:
key = self._key(session_id)
turns = self._r.incr(f"{key}:turns")
self._r.expire(f"{key}:turns", self._ttl)
return int(turns)

def set_last_meta(self, session_id: str, meta: dict) -> None:
self._r.set(f"{self._key(session_id)}:lastmeta",
json.dumps(meta), ex=self._ttl)

Wire it in by constructing the app yourself - the store lives inside create_app, so for custom storage either fork the factory function into your codebase (it's short and dependency-free apart from FastAPI itself) or keep one replica dedicated to the trap. The redis extra installs the client: pip install "tokentrap[redis]".

For rate limiting at real scale, don't move the limiter - front the service with your edge provider's rate-limiting rule and let the built-in limiter be the second line.

Health checks and operations​

GET /api/healthz (and /healthz) returns {"ok": true, "service": "tokentrap", "mode": "level-2-backend"}. Logs stream to stdout as single-line JSON - pair them with any JSON log shipper. Webhook delivery runs on a daemon thread with a four-second timeout and swallows all failures; serving never depends on logging succeeding.

When this beats the Worker​

NeedChoose
Zero servers, free tier, global edgeCloudflare Worker (level 1)
Existing FastAPI estate, internal deploymentPython backend
OpenAI bait endpointPython backend (it ships there)
Real-LLM turn-0 dressingPython backend (llm extra)
Custom session storage you controlEither - both expose small interfaces

And remember the upgrade path runs backwards too: if you start here and later want the edge, point apiEndpoint at a worker URL. Same protocol, same engine, byte-identical payloads.