JEV + LAYA | Governed AI Model Routing for Claude Code by Supreme DreamJEV + LAYA | Governed AI Model Routing for Claude Code by Supreme Dream

JEV + LAYA | Governed AI Model Routing for Claude Code

Supreme Dream

Supreme Dream

Running every turn of an AI coding session on the biggest model is the easy default and the expensive one. LAYA Code Router makes that choice per turn, on your own machine, under rules you can read.
This is part of the studio's AI systems work, alongside JEV FastLoop. Jev and Laya are different models from different teams: Jev is TypeSafe's hosted decision model, Laya is a separate open-source local model. JEV FastLoop uses Jev to route skills. LAYA Code Router uses Laya to route models. Both share one instinct: keep judgment out of the expensive path.

The brief

Cut the cost of long Claude Code sessions without a new account, a new API key or a cloud service reading the prompts. It had to run on an ordinary Claude plan, fail safe, and show its working.
How it fits together: the session, a loopback proxy, the local LAYA sidecar, then the request goes out on the chosen model
How it fits together: the session, a loopback proxy, the local LAYA sidecar, then the request goes out on the chosen model

The idea

Let a small local model score, and let plain code decide. LAYA scores each fresh turn on task complexity, reasoning required, tool complexity and one judgment question. A short policy file turns those scores into a tier. Hard work goes to Opus, ordinary coding to Sonnet, lookups to Haiku. If you type "use opus", you win. Design work never drops to Haiku.
The policy: the score cuts for the Balanced preset and the rules that sit on top of them
The policy: the score cuts for the Balanced preset and the rules that sit on top of them
Presets move the cuts, not the rules. On the project's 58 labelled prompts, Save most sends 24 to Opus, Balanced sends 31 and Careful sends 39, and in none of them does hard work land on Haiku.
Opus share by preset on the 58 labelled prompts: 24, 31 and 39
Opus share by preset on the 58 labelled prompts: 24, 31 and 39

How it is built

A launcher starts Claude Code pointed at a loopback proxy. The proxy forwards Claude Code's own headers untouched and asks a local Python sidecar for a decision. If the routing model fails, is slow or is not loaded, the turn keeps the model it already has, and no prompt waits more than 15 seconds for a decision. A macOS menu-bar app shows plan usage, the last decision and why, and an estimated saving against an all-Opus baseline, clearly labelled as an estimate.
By default a session only moves up a tier, never down, because each model keeps its own prompt cache and a downgrade starts cold. The router's own log shows why that rule matters: across the last 2,000 requests that went through the router, 96.8% of input tokens were cache reads.
From the router's log: 96.8% of input tokens were cache reads across 2,000 requests in 46 sessions
From the router's log: 96.8% of input tokens were cache reads across 2,000 requests in 46 sessions
We ran the full test suite on 2026-10-03: 747 tests, 747 passing, none skipped. Later the same day, effort routing by Anthropic's documented defaults landed and the suite grew to 790, all passing.
The real test run: npm test on 2026-10-03, 747 of 747 passing
The real test run: npm test on 2026-10-03, 747 of 747 passing
The code is open source under MIT: https://github.com/SupremeDreamZ/laya-code-router

What a client gets

AI systems that are governed, not hopeful. For a team putting models into production, we design the decision layer: cheap signals first, a clear policy in charge, a log you can audit, and a fallback that keeps work moving when a model misbehaves.

Work with us

If your AI costs are growing faster than your output, tell us how your team uses models today.
Like this project

Posted Sep 26, 2026

AI systems work: a local classifier routes each Claude Code turn to Haiku, Sonnet or Opus under a fixed policy. Open source, 790 tests passing.