⚡ Intelligent Enterprise AI Gateway · Built on LiteLLM

Route Smarter. Protect Data.
Right-Size Every LLM Request.

Flit unifies 100+ model providers behind a single OpenAI-compatible interface, automatically directs prompts to the optimal cost-performance tier, enforces enterprise DLP guardrails, and provides team-level governance.

Gateway Stream Interactive Simulator ↗ Data Plane Trace
$ POST /v1/chat/completions · Authorization: Bearer your-virtual-key
> Virtual Model Requested: "aquila" (Auto-Classification)
> Prompt: "Format this customer address list into strict GeoJSON format"
✓ Guardrail Check: Presidio DLP scanned entities · Zero PII Exfiltrated
✓ Classifier Engine: Regex Rule matched 'Formatting/GeoJSON' · Evaluated in 0.8ms
✓ Routing Decision: Routed to aquila-fast (LiteLLM deployment: gpt-4o-mini)
★ Outcome: Optimized Tier Match · Latency: 210ms · Zero PII Exposure · Spend logged to #finance
Dynamic
Multi-Tier Cost Optimization
< 2 ms
In-Memory Classifier Overhead
MCP & A2A
Protocols & Agent Governance
Zero
Inference-Path Database Queries

The Foundation of Enterprise AI

Explore why organizations adopt Flit, what capabilities it delivers out of the box, and how its decoupled architecture guarantees extreme reliability.

Why Flit?

Uncontrolled model sprawl leads to massive budget waste, security blindspots, and vendor lock-in. Flit eliminates over-provisioning by directing work to the right model tier automatically.

Explore the Problem & ROI →

What is Flit?

A unified gateway combining smart model tiering (aquila-fast, smart, power), native MCP protocol routing, A2A autonomous agents, enterprise Presidio DLP guardrails, team-scoped virtual keys, and real-time usage analytics.

See Product Capabilities →

How It Works

Complete data plane and control plane separation. LiteLLM handles high-throughput inference; workers compile routing state in memory with zero database queries on hot paths.

Deep Dive into Architecture →

The Three Public Tiers

Developers simply request aquila or one of three transparent tiers. Flit takes care of health, failover, and exact deployment selection.

aquila-fast Low Cost & Latency

Optimized for lightweight processing, formatting, classification, extraction, and high-frequency chatbot interactions.

  • Target Models GPT-4o-mini, Haiku 3.5
  • Latency < 300 ms
  • Cost Profile ~ $0.15 - $0.30 / MTok
  • Failover Target Multi-region replica
View Fast Tier Specs
aquila-power Frontier Reasoning

Reserved for complex math, formal logic proofs, deeply nested architectures, and multi-step agentic planning.

  • Target Models o1, DeepSeek-R1
  • Latency Variable (Chain of thought)
  • Cost Profile Frontier rates
  • Failover Target Dedicated fallback pool
View Power Tier Specs

Drop-in OpenAI Compatibility

Switch your base URL and API key. Existing code, SDKs, and LangChain/LlamaIndex agents work without modification.

from openai import OpenAI

# Simply point client to your Flit gateway endpoint
client = OpenAI(
    base_url="http://gateway.yourdomain.com:4000/v1",
    api_key="flit-team-scoped-virtual-key"
)

# Use virtual model selector "aquila" for automatic classification
response = client.chat.completions.create(
    model="aquila",
    messages=[
        {"role": "user", "content": "Extract the line items and totals from this invoice..."}
    ]
)

print(f"Model selected: {response.model}")
print(response.choices[0].message.content)

Ready to take control of your AI infrastructure?

Deploy Flit with Docker Compose in under five minutes. No credit card, no proprietary lock-in, open source foundation.

Read Quickstart Guide Calculate Your Savings