Home Why What How Simulator Documentation Contact Us hello@rytqlk.com
Get Started
Flit / Docs / Model Tiers & Routing

Model Tiers & Routing

Flit shields applications from model version churn by introducing a stable, intuitive vocabulary of tiers backed by intelligent runtime classifiers.

1. Public Model Vocabulary

Clients can request any of the following standard model identifiers in their request payload:

Model Identifier Classification Behavior Target Provider Deployments
aquila Automatic Classification: Evaluated by Flit's in-memory Regex/LLM classifier. Dynamically assigned to aquila-fast, aquila-smart, or aquila-power.
aquila-fast Direct Tier Routing: Sub-second response times and minimal token costs. Deployments such as gpt-4o-mini, claude-3-5-haiku, gemini-2.0-flash, llama-3.1-8b-instant.
aquila-smart Direct Tier Routing: Balanced enterprise reasoning and complex synthesis. Deployments such as gpt-4o, claude-3-5-sonnet.
aquila-power Direct Tier Routing: Deep analytical compute, STEM proofs, and heavy coding. Deployments such as o1, deepseek-r1.
<exact_deployment_id> Explicit Pinning: Directly selects a specific LiteLLM deployment. Directly to assigned deployment (intentionally bypasses group failover).

2. Classifier Engine Modes

Administrators configure the active classifier workspace under the Flit Web Control Plane (/admin/classifiers). Three modes are supported:

Mode A: Regex Classifier

Executes an ordered list of regular expression patterns against the input prompt in worker memory with zero network or database I/O. If a pattern matches, the assigned tier is returned immediately. If no rule matches, it falls back to the configured default tier.

Sample Regex Configuration (JSON)
[
  {
    "id": "rule-translation",
    "name": "Multilingual Translation",
    "pattern": "^(?i)(translate|translation)\\b",
    "target_tier": "aquila-fast",
    "priority": 10
  },
  {
    "id": "rule-deep-coding",
    "name": "Architecture & Algorithms",
    "pattern": "(?i)(distributed algorithm|consensus|formal proof|compiler)",
    "target_tier": "aquila-power",
    "priority": 20
  }
]

Mode B: Classifier LLM

Evaluates incoming prompts against a structured system prompt template using candidate deployments registered in the protected aquila-internal-classifier group. The gateway dispatches the classification self-call using an internal system credential and sets the X-Aquila-Internal-Purpose header to prevent recursive loops.

Mode C: Hybrid Engine (Recommended)

Evaluates Regex rules first in sub-millisecond memory. If any rule triggers, the tier is assigned instantly. Only upon NO_REGEX_RULE_MATCH does it seamlessly transition to the Classifier LLM for nuanced semantic classification.

3. Supported Providers & Verified Access Methods

Flit provides a curated provider catalog with five verified, production-tested access methods:

Provider Family Access Method LiteLLM Route Prefix Credential Fields
OpenAI openai_api openai/ Required api_key; optional api_base
Anthropic anthropic_api anthropic/ Required api_key
Google gemini_api gemini/ Required Google AI Studio api_key
Groq groq_api groq/ Required api_key (supports slash-delimited models like openai/gpt-oss-120b)
Ollama ollama_api ollama/ Optional api_base (e.g., http://host.docker.internal:11434); no key needed

Note: Microsoft Azure and AWS Bedrock access methods are in active development and slated for the next release.

4. Same-Group Failover & Health Monitoring

Within each logical tier group (e.g., aquila-smart), administrators can register multiple deployments across different providers. LiteLLM dynamically monitors deployment health:

  • If an upstream provider returns HTTP 429 or 5xx, the request is automatically retried across alternative healthy deployments in the same group.
  • Failovers remain completely transparent to the client application with zero manual configuration.