Route Smarter. Protect Data.
Right-Size Every LLM Request.
Flit unifies 100+ model providers behind a single OpenAI-compatible interface, automatically directs prompts to the optimal cost-performance tier, enforces enterprise DLP guardrails, and provides team-level governance.
The Foundation of Enterprise AI
Explore why organizations adopt Flit, what capabilities it delivers out of the box, and how its decoupled architecture guarantees extreme reliability.
Why Flit?
Uncontrolled model sprawl leads to massive budget waste, security blindspots, and vendor lock-in. Flit eliminates over-provisioning by directing work to the right model tier automatically.
What is Flit?
A unified gateway combining smart model tiering (aquila-fast, smart, power), native MCP protocol routing, A2A autonomous agents, enterprise Presidio DLP guardrails, team-scoped virtual keys, and real-time usage analytics.
How It Works
Complete data plane and control plane separation. LiteLLM handles high-throughput inference; workers compile routing state in memory with zero database queries on hot paths.
The Three Public Tiers
Developers simply request aquila or one of three transparent tiers. Flit takes care of health, failover, and exact deployment selection.
Optimized for lightweight processing, formatting, classification, extraction, and high-frequency chatbot interactions.
- Target Models GPT-4o-mini, Haiku 3.5
- Latency < 300 ms
- Cost Profile ~ $0.15 - $0.30 / MTok
- Failover Target Multi-region replica
The versatile workhorse for multi-turn synthesis, coding, document reasoning, and nuanced natural language responses.
- Target Models GPT-4o, Sonnet 3.5
- Latency < 850 ms
- Cost Profile ~ $2.50 - $3.00 / MTok
- Failover Target Cross-provider backup
Reserved for complex math, formal logic proofs, deeply nested architectures, and multi-step agentic planning.
- Target Models o1, DeepSeek-R1
- Latency Variable (Chain of thought)
- Cost Profile Frontier rates
- Failover Target Dedicated fallback pool
Drop-in OpenAI Compatibility
Switch your base URL and API key. Existing code, SDKs, and LangChain/LlamaIndex agents work without modification.
from openai import OpenAI
# Simply point client to your Flit gateway endpoint
client = OpenAI(
base_url="http://gateway.yourdomain.com:4000/v1",
api_key="flit-team-scoped-virtual-key"
)
# Use virtual model selector "aquila" for automatic classification
response = client.chat.completions.create(
model="aquila",
messages=[
{"role": "user", "content": "Extract the line items and totals from this invoice..."}
]
)
print(f"Model selected: {response.model}")
print(response.choices[0].message.content)