Hybrid Inference — Local + Cloud

Your AI bill is a
routing problem.

Token Station's smart router sends simple tasks to local models at near-zero cost, and hard tasks to GPT-5.5 and Claude Opus — then trains itself on your data every 60 days to get smarter.

No credit card required · One OpenAI-compatible endpoint · Works with your existing stack

YOUR APP API Request SMART ROUTER Token Station complexity classifier ● live simple complex LOCAL MODELS Llama · Mistral ≈$0.000 per call CLOUD MODELS GPT-5.5 · Claude high-complexity tasks 73% avg. cost reduction Token Station composite SELF-EVOLVING RL fine-tuning trained on YOUR data
%
Cost Reduction
20×
Cheaper vs. Claude Opus Only
%
Fewer Large-Model Calls
+
Supported Models
d
To First Custom Router

You're routing every request to the most expensive model by default

Most teams route everything to GPT-5.5 because it's the safest choice — not because every task needs it.

💸

Paying GPT-5.5 prices for tasks Llama handles fine

Summarization, Q&A, formatting, classification — 60-70% of your API calls don't need frontier reasoning. You're paying for a jet engine to fly across the street.

$0.011/1K tokens (GPT-5.5) → $0.0001/1K locally
📉

Static routing rules go stale as your workload shifts

You write routing rules in Week 1, then your product changes. The rules stay. A fintech's "simple" isn't a law firm's "simple" — and neither is yours six months from now.

LiteLLM · Portkey · OpenRouter: all static rules
🖥️

Local models sit idle while your cloud bill grows

You bought the compute, deployed the models, and now 90% of your traffic bypasses them entirely. Your local GPU cluster is a sunk cost waiting to be unlocked.

Avg. local GPU utilization: <15% in AI teams

Route. Learn. Optimize.

Three phases that compound over time — starting with rules on day one, graduating to a self-trained classifier within 60 days.

Step 01
🔀

Route by Rules

On day one, configure routing with user-defined rules: "tasks under 500 tokens → Haiku," "use GPT-5.5 above confidence threshold." No redeployment needed to change rules.

route(task)
→ complexity_score: 0.23
→ model: llama-3-8b
→ cost: $0.0001
Step 02
🧠

Learn from Your Data

After 60 days, Token Station trains a custom complexity classifier on your actual request history. It learns what "simple" and "complex" mean for your specific workloads — not a generic benchmark.

training_data: 847K requests
classifier: fine-tuned on your corpus
accuracy: +31% vs static rules
Step 03
⚡

Optimize Forever

Reinforcement learning refines the router using real quality signals — user ratings, downstream task success, latency. More usage means better accuracy, which means lower costs. Automatically.

rl_reward: quality × speed × cost
router_version: v4 (auto-updated)
savings_trend: ↑ 8% / 30 days

Built for ML infrastructure engineers

Not another wrapper. A production-grade inference layer with learning capabilities no other product ships.

Hybrid Inference

One endpoint. Two cost tiers. Zero code changes.

Drop Token Station in front of your existing OpenAI calls. Your team doesn't change a line of code — the router decides whether each request goes to Llama 3 on your Olares Mini or GPT-5.5 in the cloud.

  • OpenAI-compatible API — /v1/chat/completions works as-is
  • Local models: Llama, Mistral, Phi, Qwen on your own hardware
  • Data never leaves your network for locally-routed calls
  • Automatic fallback when local model confidence is below threshold
  • Latency caps, cost floors, quality thresholds — all configurable
AI inference routing diagram
Routing Decision
73% → Local
Self-Evolving Router

The router trains itself on your workload. No competitor does this.

LiteLLM, Portkey, and OpenRouter all use static rules or generic benchmarks. Token Station trains a custom classifier every 60 days using reinforcement learning on your company's actual usage history.

  • Fine-tunes on 60+ days of your actual request history
  • RL signal: quality ratings, downstream task success, latency
  • A law firm's "simple" ≠ a fintech's "simple" — it learns yours
  • Router improves automatically: more usage = better accuracy
🔬

LiteLLM, Portkey, and OpenRouter all use static routing rules — none train a model on your data. This is the moat.

Machine learning router training visualization
Router Version
v4 · Auto-trained
Local Integration

Your hardware. Your data. Near-zero marginal cost.

Token Station runs Llama, Mistral, Phi, and Qwen locally on your Olares Mini or any server you already own. Every locally-routed call costs you compute cycles, not API dollars.

  • Runs on Olares Mini — a compact personal server for self-hosted AI
  • Local inference = zero API cost per routed query
  • Sensitive data never leaves your network for local calls
  • Confidence threshold fallback — always cloud-backed for quality
  • No GPU required — optimized models run on CPU + RAM
Local AI model server hardware
Local API Cost
$0.0001 / 1K tokens
Cost Controls

Set per-employee budgets. Stop runaway AI spend before it happens.

Routing cuts costs at the infrastructure level. Quota management cuts them at the human level. Assign monthly token budgets per employee or per team — the gateway enforces limits automatically, so no single power-user can blow the quarter's budget.

  • Monthly token budgets per employee, team, or department
  • Live usage dashboard — who's spending what, on which models
  • Soft warnings at 80% quota; hard cutoff at 100%
  • When cloud budget is hit, automatically fall back to local models
  • Manager view: department-level spend rolled up in one place
  • Automatic quota resets on billing period rollover
Quota Dashboard · April 2026
Engineering · 12 seats 91% ⚠
455K / 500K tokens · routing to local ↓
Marketing · 6 seats 58%
174K / 300K tokens
Design · 4 seats 31%
62K / 200K tokens
Total · April
691K / 1M tokens
Routing Savings
$847 saved

The numbers are peer-reviewed

Hybrid inference routing isn't a startup pitch — it's a published research area from UC Berkeley, Microsoft, and Anyscale. Token Station ships it as a product.

0%
Token cost reduction via hybrid routing
Token Station Composite
2×
Cost savings vs. always using GPT-4, zero quality loss
RouteLLM · UC Berkeley / Anyscale
0%
Fewer large-model calls via difficulty-based routing
Hybrid LLM · Microsoft Research
30+
Unseen models dynamically routed without retraining
UniRoute · ICLR 2026
ICLR 2025 · UC Berkeley / Anyscale

RouteLLM: Learning to Route LLMs with Preference Data

Demonstrates that routing between strong and weak LLMs can achieve 2× or more cost savings on real-world benchmarks while maintaining response quality.

↳ 2× savings · zero quality regression
ICLR 2024 · Microsoft Research

Hybrid LLM: Cost-Efficient and Quality-Aware Query Routing

Introduces difficulty-based routing between local and cloud models, showing 40% fewer large-model calls with on-par downstream task performance.

↳ 40% fewer GPT-4 calls · same quality
ICLR 2026

UniRoute: Scalable Dynamic Routing Across Unseen Models

Demonstrates dynamic routing across 30+ previously unseen models without retraining the router, proving generalization at production scale.

↳ 30+ models · zero retraining needed

Scales with your team, not your usage

Flat per-seat pricing means the more you route, the better your unit economics get.

On-Prem
$10 / employee / month
Bring your own API keys. Token Station runs on your infrastructure.
  • Bring your own API keys (OpenAI, Anthropic, Google)
  • Local model support — Llama, Mistral, Phi, Qwen
  • Self-evolving router after 60 days
  • All routing rules and cost thresholds
  • Per-employee token quotas and usage alerts
  • Data stays on your network
  • Slack + email support
No credit card required

Technical questions, answered

Does routing add latency to every request? +
Token Station's complexity classifier runs in under 5ms — far below the P99 latency of any LLM call. For local inference, you eliminate round-trip network latency entirely, which often makes locally-routed requests faster than cloud calls. For cloud-routed requests, the routing overhead is negligible. We publish latency benchmarks in the pilot dashboard.
Will routing to smaller models hurt quality? +
That's the whole problem we solved. The router only sends tasks to local models when its confidence score exceeds a threshold you configure (default: 0.85). If the classifier isn't confident, it falls back to cloud. The RouteLLM paper (ICLR 2025) demonstrates 2× cost savings with zero measurable quality regression on real-world benchmarks — and our router improves on that baseline using your specific data.
Which local models are supported? +
Token Station currently supports Llama 3 (8B, 70B), Mistral 7B, Phi-3 (mini, small, medium), and Qwen 2.5. We add new models quarterly. All models run via the Olares runtime — no manual CUDA setup required. You can also plug in any model that exposes an OpenAI-compatible endpoint.
What data does the self-evolving router train on? +
The router trains only on metadata about your requests — token counts, task categories, confidence scores, and downstream quality signals you opt into. No request content (prompts or completions) is ever used for training. The trained classifier stays in your environment; it does not leave your network. Training runs on-premise using Olares compute.
How is this different from LiteLLM or Portkey? +
LiteLLM and Portkey are excellent load-balancers and proxies — they route by rules you define. Token Station adds a learning layer on top: the complexity classifier trains on your actual usage and improves over time. Static rules decay as your product evolves. A self-trained router doesn't. That's the fundamental difference.
Can I set limits so one team doesn't blow the whole budget? +
Yes — quota management is a core feature. You assign monthly token budgets per employee, per team, or per department from the admin dashboard. The gateway enforces limits in real time: a soft warning fires at 80% usage, and a hard cutoff triggers at 100%. Once a cloud budget is exhausted, the router automatically shifts that employee's traffic to local models — they stay productive, but your cloud bill doesn't spike. Quotas reset automatically at the start of each billing period.

Start cutting your
AI costs today.

30-day free pilot. One OpenAI-compatible endpoint. Your first custom router trained in 60 days.

No credit card · No sales call · Works with GPT-5.5, Claude Opus, Gemini, and 30+ models