What is Flit?
The Enterprise Gateway Suite
Flit provides the complete control plane and data plane necessary to operate generative AI at enterprise scale with confidence, security, and predictability.
Unified Multi-Model Gateway
Flit standardizes major commercial and open-weights LLMs into a strict, unified OpenAI-compatible protocol. Connect OpenAI, Anthropic Claude, Google Gemini (AI Studio), Groq, and Ollama, alongside native Model Context Protocol (MCP) servers and external Agent-to-Agent (A2A) entities.
- ✓ Standard endpoints:
/v1/chat/completions,/v1/responses,/v1/models - ✓ Native MCP Streamable HTTP & SSE data plane at
/mcp - ✓ A2A Agent discovery and JSON-RPC dispatch at
/a2a/*and/v1/agents - ✓ Transparent failover across alternative deployments in the same tier
- ✓ Private upstream credential vault with zero secret leakage to browsers
[
{ "provider": "openai", "access_method": "openai_api", "models": ["gpt-4o", "gpt-4o-mini", "o1"] },
{ "provider": "anthropic", "access_method": "anthropic_api", "models": ["claude-3-5-sonnet", "claude-3-5-haiku"] },
{ "provider": "google", "access_method": "gemini_api", "models": ["gemini-2.0-flash", "gemini-1.5-pro"] },
{ "provider": "groq", "access_method": "groq_api", "models": ["llama-3.1-8b-instant", "openai/gpt-oss-120b"] },
{ "provider": "ollama", "access_method": "ollama_api", "models": ["llama3.1:latest", "mistral:latest"] }
]
// Azure OpenAI & AWS Bedrock coming next
// Client requests virtual model:
{
"model": "aquila", // Automatic classification
"messages": [...]
}
// Or select an explicit performance tier:
// - "aquila-fast" (High speed, low cost)
// - "aquila-smart" (Balanced reasoning)
// - "aquila-power" (Frontier reasoning)
// - "exact_id" (Deliberate model pin)
Small, Stable Model Vocabulary
Instead of forcing hundreds of engineers to manage individual model version IDs, Flit offers a clean, four-term vocabulary. Developers write against stable tiers while administrators swap underlying deployments behind the scenes with zero client downtime.
- ✓
aquila: Dynamic semantic auto-classification - ✓
aquila-fast: Low-cost text work ($0.15 - $0.30 / MTok) - ✓
aquila-smart: Balanced enterprise workhorse - ✓
aquila-power: Highest-capability frontier reasoning
Flexible Classifier Workbenches
Three interchangeable classification modes suited for any latency and accuracy requirement.
Regex Classifier
Evaluates incoming prompts against pre-compiled regular expression patterns. Operates in <1 millisecond with zero external inference calls.
Classifier LLM
Uses a fast, low-cost classifier prompt to analyze request intent, complexity, and safety parameters before selecting the target tier.
Hybrid Engine
Combines instant Regex pattern matching with automatic Classifier LLM fallback when no deterministic rule triggers.
Enterprise DLP & Guardrails
Integrated with Microsoft Presidio Analyzer to inspect incoming user prompts and outgoing model completions for sensitive data entities.
- 🛡️ Automatic PII detection: SSNs, credit cards, emails, IP addresses
- 🛡️ Configurable redaction, masking, or immediate blocking
- 🛡️ Prompt injection defense and system override shielding
- 🛡️ Comprehensive audit logs for compliance reviews
Multi-Tenant Teams & Budgets
Hierarchical team organization allows enterprises to partition models, quotas, and virtual keys across business departments.
- 👥 Microsoft Entra ID (Azure AD) SSO and Local Auth
- 👥 Role-Based Access Control: Global Admin, Team Admin, Member
- 👥 Team-scoped Virtual Keys with hard TPM/RPM rate limits
- 👥 Monthly spend caps with automatic circuit breaking