⚙️ Architectural Deep Dive

How Flit Works
Engineered for Extreme Reliability

Flit enforces strict separation between data and control planes, eliminating database latency from the inference hot path and ensuring high-availability routing across commercial and private models.

Data Plane vs. Control Plane Separation

Inference queries pass directly from Nginx to the LiteLLM gateway. They never traverse the web application layer.

Inference & Protocol Data Plane Sub-millisecond

High-Throughput Gateway

Handles /v1/chat/completions, /v1/responses, /v1/models, /mcp, and /a2a.

  • Worker Memory Cache: In-memory compiled classifier rules & local routing tables
  • LiteLLM Core: Streaming SSE, provider adapters, load balancing & failover
  • Zero Database Queries: Hot paths execute completely in worker process memory
  • Protocol Gateways: Streamable HTTP / SSE for MCP and JSON-RPC for A2A
Admin Control Plane Authoritative State

Web Application & Policy Core

Handles browser identity, Teams, Virtual Keys, MCP servers, A2A agents, and Classifiers.

  • Identity & Auth: Entra ID SSO, Local Admin, Fernet encrypted cookies
  • PostgreSQL SSOT: Immutable audit trail in admin_audit_log & singleton classifier row
  • Zero-IO State Sync: 30s background hydration into worker memory; no hot-path DB I/O
  • Isolated Network: Master keys and admin routes remain strictly private

End-to-End Request & Response Flow

Watch a query journey in real time: starting from the browser or API client, passing through ingress scrubbing and zero-IO in-memory routing, sanitizing with DLP, executing at the LLM, and streaming back with asynchronous telemetry.

💻
Browser / API

User chat or automated backend client

Virtual Key Auth
🌐
Nginx Ingress

Port 4000 edge proxy & header scrubber

TLS & Rate Limit
⚡
Flit Gateway

In-Memory Classifier & Presidio DLP

< 2ms Memory Rule DLP Mask / Redact
🧠
LLM Provider

OpenAI, Claude, Bedrock, or Private vLLM

SSE Token Stream
FLOW PHASE: OUTBOUND REQUEST
LATENCY: 0.0 ms
// Step 1: Application / Browser Client dispatches request
POST http://api.flit.internal:4000/v1/chat/completions
Authorization: Bearer your-virtual-key
Content-Type: application/json

{
  "model": "aquila", // Auto-classification requested
  "messages": [
    {"role": "user", "content": "Extract customer tax ID and format as JSON: SSN 000-12-3456"}
  ]
}

The Path of an Inference Query

Follow a request from client dispatch to upstream completion and asynchronous telemetry logging.

1

Ingress Ingestion

Nginx receives request on host port 4000, strips sensitive internal headers, validates TLS, and proxies directly to the LiteLLM gateway worker pool.

2

Key & Quota Check

Worker verifies the virtual key, confirms team permissions, checks budget caps, and applies token-per-minute (TPM) rate limits in memory.

3

Zero-IO Classifier

If model is aquila, the compiled in-memory regex or hybrid classifier resolves the prompt to aquila-fast, smart, or power in <1.5ms.

4

Provider Execution

LiteLLM routes prompt to healthy provider deployments in the selected tier, managing retries, latency balancing, and automatic failovers.

Durable State & Optimistic Concurrency

Flit guarantees that administrative mutations (such as changing classifier rules or updating team quotas) never expose inconsistent state or corrupt runtime workers.

1. Optimistic Checksum

Mutations require the current workspace checksum and CSRF token. Out-of-sync edits return HTTP 409 Conflict, preventing accidental administrative overwrites.

2. PostgreSQL Commit

Authoritative state is committed to PostgreSQL in a single current configuration row alongside an immutable event record in admin_audit_log.

3. Resilient In-Memory Compilation

Gateway workers poll PostgreSQL every 30s in the background and compile rules directly into worker memory. If a poll fails, workers gracefully retain valid in-memory state.

Deployment Options

From local developer workstation to high-availability multi-cloud clusters.

Developer Sandbox

Local Docker Compose

Spin up LiteLLM, PostgreSQL, Redis, and Flit Web in a single command. Perfect for local prototyping, unit testing, and team evaluations.

docker compose up -d
Single-Cloud Production

Container Services

Deploy to AWS ECS, Azure Container Apps, or Google Cloud Run with managed PostgreSQL (RDS/Cloud SQL) and managed Redis (ElastiCache/MemoryStore).

Includes health checks, autoscaling, and TLS termination.
Enterprise Scale

Kubernetes & Multi-Region

Production Helm charts with horizontal pod autoscalers (HPA), dedicated ingress controllers, distributed Redis cluster, and cross-region failover.

99.99% availability SLA with zero-downtime rolling upgrades.

Start building with Flit today

Review our comprehensive documentation or explore the interactive classifier simulator.

Read Getting Started Launch Simulator