How routing works
This is the idea at the centre of FRIDAY. Cloud agents route tools well because the model is huge. Strip the cloud away and a small on-device model can't be trusted to pick the right tool — so FRIDAY puts a deterministic router in front of it.
Why local-first agents are hard
An agent harness is the loop that reads your request, decides which tool to call, and fills that tool's arguments. Modern harnesses lean entirely on a frontier model to do this. Run the same harness on a 0.8B model that fits on a laptop and it breaks down: it misreads intent, picks the wrong tool, or — most dangerously — fabricates a successful-sounding result for an action it never performed.
“Brightness set to 60.” Nothing changed. That confident lie is what FRIDAY is built to prevent.The bridge: a deterministic router
FRIDAY adds a routing layer that doesn't depend on the model's intelligence for the common cases. It combines three deterministic techniques, each defending a confidence band so a near-miss asks rather than guesses:
- Regex intent matching — hand-written parsers turn known phrasings straight into a tool call.
- Similarity — fuzzy (rapidfuzz) and embedding (cosine) scoring catch paraphrases and speech-to-text slips.
- Confidence rules — explicit thresholds decide between dispatch, “did you mean…?”, and falling through.
The six layers, cheapest first
PlannerEngine.plan() walks the chain and returns at the first layer that produces a confident plan. A regex hit at L1 (confidence ≥ 0.90) short-circuits everything else — the model is never consulted, so it never gets to invent a result.
pre-checks 1. pending online confirmation 2. active workflow resume
──────────────────────────────────────────────────────────────────────
L1 IntentRecognizer deterministic regex parsers (first-match wins) conf 1.0 / 0.9
L2 RouteScorer alias / pattern / context-term score (min 80) "score"
L2a Learned dispatch a phrasing confirmed promote_after times "learned"
L2b LexicalRouter rapidfuzz token_set_ratio (catches STT / typos) fuzzy
L3 EmbeddingRouter cosine over capability embeddings cosine band
L4 QwenPlanner local 4B planner (LLM tool / plan synthesis) generative
L5 llm_chat conversational fallback chatConfidence bands
The fuzzy layers never dispatch on a bare best-match; each defends a band. These are the defaults (all tunable under routing.* — see Configuration):
Intent fast-path HIGH_THRESHOLD 0.90 at/above → bypass planner, build plan from regex
Intent fast-path MEDIUM_THRESHOLD 0.50 below HIGH but here → keep candidate, still consult planner
Embedding dispatch_threshold 0.62 cosine ≥ this → auto-dispatch matched capability
Embedding confirm_low 0.50 in [confirm_low, dispatch) → ask "did you mean…?"
Embedding tie_epsilon 0.05 two candidates this close → disambiguate
Lexical lexical_threshold 88 rapidfuzz score must clear this to fire
Lexical lexical_margin 6 …and beat the runner-up by this margin (never poaches)
Learned promote_after 3 a phrasing confirmed this many times → deterministicpromote_after counter. FRIDAY gets more confident about your phrasing the more you use it.The tradeoff — read this
Determinism is a deal, not magic. A regex router only fires for phrasings it was taught. Give FRIDAY a complex or novel command it has never seen and the deterministic layers won't catch it — it falls back to the local 4B planner, which is far weaker than a frontier model and may not get it right.
This is the honest cost of local-first: for every tool we build, we also teach FRIDAY the words that drive it. The payoff is reliability, privacy, and zero cloud dependence for everything it does know — plus a system that learns new phrasings as you use it. When you add a capability, you add its intent pattern too; that's the subject of Add a new tool.