Model Tiers & Routing
Flit shields applications from model version churn by introducing a stable, intuitive vocabulary of tiers backed by intelligent runtime classifiers.
1. Public Model Vocabulary
Clients can request any of the following standard model identifiers in their request payload:
| Model Identifier | Classification Behavior | Target Provider Deployments |
|---|---|---|
aquila |
Automatic Classification: Evaluated by Flit's in-memory Regex/LLM classifier. | Dynamically assigned to aquila-fast, aquila-smart, or aquila-power. |
aquila-fast |
Direct Tier Routing: Sub-second response times and minimal token costs. | Deployments such as gpt-4o-mini, claude-3-5-haiku, gemini-2.0-flash, llama-3.1-8b-instant. |
aquila-smart |
Direct Tier Routing: Balanced enterprise reasoning and complex synthesis. | Deployments such as gpt-4o, claude-3-5-sonnet. |
aquila-power |
Direct Tier Routing: Deep analytical compute, STEM proofs, and heavy coding. | Deployments such as o1, deepseek-r1. |
<exact_deployment_id> |
Explicit Pinning: Directly selects a specific LiteLLM deployment. | Directly to assigned deployment (intentionally bypasses group failover). |
2. Classifier Engine Modes
Administrators configure the active classifier workspace under the Flit Web Control Plane (/admin/classifiers). Three modes are supported:
Mode A: Regex Classifier
Executes an ordered list of regular expression patterns against the input prompt in worker memory with zero network or database I/O. If a pattern matches, the assigned tier is returned immediately. If no rule matches, it falls back to the configured default tier.
[
{
"id": "rule-translation",
"name": "Multilingual Translation",
"pattern": "^(?i)(translate|translation)\\b",
"target_tier": "aquila-fast",
"priority": 10
},
{
"id": "rule-deep-coding",
"name": "Architecture & Algorithms",
"pattern": "(?i)(distributed algorithm|consensus|formal proof|compiler)",
"target_tier": "aquila-power",
"priority": 20
}
]
Mode B: Classifier LLM
Evaluates incoming prompts against a structured system prompt template using candidate deployments registered in the protected aquila-internal-classifier group. The gateway dispatches the classification self-call using an internal system credential and sets the X-Aquila-Internal-Purpose header to prevent recursive loops.
Mode C: Hybrid Engine (Recommended)
Evaluates Regex rules first in sub-millisecond memory. If any rule triggers, the tier is assigned instantly. Only upon NO_REGEX_RULE_MATCH does it seamlessly transition to the Classifier LLM for nuanced semantic classification.
3. Supported Providers & Verified Access Methods
Flit provides a curated provider catalog with five verified, production-tested access methods:
| Provider Family | Access Method | LiteLLM Route Prefix | Credential Fields |
|---|---|---|---|
| OpenAI | openai_api |
openai/ |
Required api_key; optional api_base |
| Anthropic | anthropic_api |
anthropic/ |
Required api_key |
gemini_api |
gemini/ |
Required Google AI Studio api_key |
|
| Groq | groq_api |
groq/ |
Required api_key (supports slash-delimited models like openai/gpt-oss-120b) |
| Ollama | ollama_api |
ollama/ |
Optional api_base (e.g., http://host.docker.internal:11434); no key needed |
Note: Microsoft Azure and AWS Bedrock access methods are in active development and slated for the next release.
4. Same-Group Failover & Health Monitoring
Within each logical tier group (e.g., aquila-smart), administrators can register multiple deployments across different providers. LiteLLM dynamically monitors deployment health:
- If an upstream provider returns HTTP 429 or 5xx, the request is automatically retried across alternative healthy deployments in the same group.
- Failovers remain completely transparent to the client application with zero manual configuration.