How Flit Works
Engineered for Extreme Reliability
Flit enforces strict separation between data and control planes, eliminating database latency from the inference hot path and ensuring high-availability routing across commercial and private models.
Data Plane vs. Control Plane Separation
Inference queries pass directly from Nginx to the LiteLLM gateway. They never traverse the web application layer.
High-Throughput Gateway
Handles /v1/chat/completions, /v1/responses, /v1/models, /mcp, and /a2a.
- Worker Memory Cache: In-memory compiled classifier rules & local routing tables
- LiteLLM Core: Streaming SSE, provider adapters, load balancing & failover
- Zero Database Queries: Hot paths execute completely in worker process memory
- Protocol Gateways: Streamable HTTP / SSE for MCP and JSON-RPC for A2A
Web Application & Policy Core
Handles browser identity, Teams, Virtual Keys, MCP servers, A2A agents, and Classifiers.
- Identity & Auth: Entra ID SSO, Local Admin, Fernet encrypted cookies
- PostgreSQL SSOT: Immutable audit trail in
admin_audit_log& singleton classifier row - Zero-IO State Sync: 30s background hydration into worker memory; no hot-path DB I/O
- Isolated Network: Master keys and admin routes remain strictly private
End-to-End Request & Response Flow
Watch a query journey in real time: starting from the browser or API client, passing through ingress scrubbing and zero-IO in-memory routing, sanitizing with DLP, executing at the LLM, and streaming back with asynchronous telemetry.
Browser / API
User chat or automated backend client
Nginx Ingress
Port 4000 edge proxy & header scrubber
Flit Gateway
In-Memory Classifier & Presidio DLP
LLM Provider
OpenAI, Claude, Bedrock, or Private vLLM
// Step 1: Application / Browser Client dispatches request
POST http://api.flit.internal:4000/v1/chat/completions
Authorization: Bearer your-virtual-key
Content-Type: application/json
{
"model": "aquila", // Auto-classification requested
"messages": [
{"role": "user", "content": "Extract customer tax ID and format as JSON: SSN 000-12-3456"}
]
}
The Path of an Inference Query
Follow a request from client dispatch to upstream completion and asynchronous telemetry logging.
Ingress Ingestion
Nginx receives request on host port 4000, strips sensitive internal headers, validates TLS, and proxies directly to the LiteLLM gateway worker pool.
Key & Quota Check
Worker verifies the virtual key, confirms team permissions, checks budget caps, and applies token-per-minute (TPM) rate limits in memory.
Zero-IO Classifier
If model is aquila, the compiled in-memory regex or hybrid classifier resolves the prompt to aquila-fast, smart, or power in <1.5ms.
Provider Execution
LiteLLM routes prompt to healthy provider deployments in the selected tier, managing retries, latency balancing, and automatic failovers.
Durable State & Optimistic Concurrency
Flit guarantees that administrative mutations (such as changing classifier rules or updating team quotas) never expose inconsistent state or corrupt runtime workers.
Mutations require the current workspace checksum and CSRF token. Out-of-sync edits return HTTP 409 Conflict, preventing accidental administrative overwrites.
Authoritative state is committed to PostgreSQL in a single current configuration row alongside an immutable event record in admin_audit_log.
Gateway workers poll PostgreSQL every 30s in the background and compile rules directly into worker memory. If a poll fails, workers gracefully retain valid in-memory state.
Deployment Options
From local developer workstation to high-availability multi-cloud clusters.
Local Docker Compose
Spin up LiteLLM, PostgreSQL, Redis, and Flit Web in a single command. Perfect for local prototyping, unit testing, and team evaluations.
docker compose up -d
Container Services
Deploy to AWS ECS, Azure Container Apps, or Google Cloud Run with managed PostgreSQL (RDS/Cloud SQL) and managed Redis (ElastiCache/MemoryStore).
Kubernetes & Multi-Region
Production Helm charts with horizontal pod autoscalers (HPA), dedicated ingress controllers, distributed Redis cluster, and cross-region failover.