System Architecture
Flit is designed for extreme operational throughput, enterprise resilience, and clear operational boundaries. This document details ingress routing, plane separation, in-memory classification, and durable state consistency.
1. Ingress Boundaries & Reverse Proxy
Nginx is the sole public ingress point into the system, operating on host port 4000. It enforces strict route segregation, scrubs internal topology headers, and isolates the high-throughput data plane from the administrative control plane:
| Path Pattern | Target Upstream | Data / Control Plane | Description |
|---|---|---|---|
/v1/chat/completions |
flit_gateway:4010 |
Inference Data Plane | Direct OpenAI-compatible inference path. Bypasses web application entirely. |
/v1/responses |
flit_gateway:4010 |
Inference Data Plane | Native responses endpoint for structured text and file inputs. |
/v1/models |
flit_gateway:4010/v1/aquila/models |
Inference Data Plane | Resolved to caller-filtered catalog of public tiers (aquila-fast, etc.) and permitted exact deployments. |
/mcp & /mcp/ |
flit_gateway:4010 |
MCP Data Plane | Native Model Context Protocol (Streamable HTTP & SSE) data plane authenticated via x-litellm-api-key. |
/a2a/*, /v1/a2a/*, /v1/agents |
flit_gateway:4010 |
A2A Agent Data Plane | Agent-to-Agent discovery (agent-card) and JSON-RPC execution scoped to authorized virtual keys. |
/admin/* & /* |
flit_web:4005 |
Admin Control Plane | Browser management UI, Microsoft Entra SSO callbacks, and Flit administrative REST APIs. |
X-Litellm-Model-Id, X-Litellm-Model-Api-Base, and X-Litellm-Model-Group are stripped by Nginx so internal infrastructure details are never leaked. In addition, X-Aquila-Internal-Purpose is stripped on public ingress so external callers cannot bypass classifier controls.
2. Data Plane vs Control Plane
Flit enforces a hard separation between the Inference / Protocol Data Plane and the Administrative Control Plane:
- LiteLLM Gateway (Data Plane): Owns provider protocol translation, retry logic, same-group failover, provider rate limits, raw token/spend accounting, and MCP / A2A data plane dispatch. It communicates with upstream LLM APIs and local Redis caches.
- Flit Web (Control Plane): Owns user identity, Microsoft Entra ID SSO sessions, local password management, Team & Key administration, MCP server management, external A2A agent registration, and classifier policy editing.
3. Zero-IO In-Memory Routing
Traditional gateways query a relational database or external cache on every single prompt to determine routing rules. Under high concurrency, this leads to connection pool exhaustion and unpredictable latency spikes.
Flit eliminates hot-path database I/O completely. Gateway workers hydrate the active classifier workspace from PostgreSQL into local memory in the background (default: every 30 seconds). When an inference call arrives:
- The request payload is read from the worker process memory.
- Pre-compiled regular expressions and intent rules evaluate in microseconds.
- The target model group (
aquila-fast,aquila-smart, oraquila-power) is assigned. - LiteLLM native router handles group deployment dispatch and healthy replica selection.
4. Durable State & Optimistic Concurrency
To maintain strict operational predictability and eliminate data desynchronization:
- PostgreSQL Authoritative Storage: Flit stores classifier configuration as a single current row in
aquila_core.classifier_configurationalongside immutable audit records inadmin_audit_log. Flit does not maintain complex draft queues, rollback pools, or Redis classifier snapshots. - Optimistic Concurrency Control: Administrative updates require the current workspace checksum (along with CSRF token and change summary). If another administrator saved a newer decision in the meantime, the mutation returns HTTP 409 Conflict, preventing accidental overwrites.
- Graceful Degradation: If a gateway worker encounters a temporary database error or invalid row during background polling, it continues executing with its existing compiled in-memory configuration, guaranteeing zero dropped inference requests.