The assistant that
stays on your machine.
FRIDAY listens, reasons, and acts entirely on-device. Speech-to-text, the conversational model, planning, vision, and memory all run on your hardware — it reaches the internet only when you ask. Say "Hey Friday" and it gets to work, narrating progress out loud.
git clone https://github.com/SanthoshReddy352/Friday_Linux.git
cd Friday_Linux && ./setup.sh
python main.pyCloud agents are smart because the model is smart. Take the cloud away and the harness falls apart.
The current wave of agent harnesses — the loops that read your request, pick a tool, and fill its arguments — lean entirely on a frontier model to understand intent and route correctly. Run that same harness on a model small enough to live on your laptop and it breaks: small models miss the intent, call the wrong tool, or worse —
Small models hallucinate success
Ask a 0.8B chat model to set your brightness and — with no tool actually wired — it will cheerfully reply "Brightness set to 60." Nothing happened. The model fabricated a plausible-sounding result because that's what language models do when they can't really act.
FRIDAY bridges the gap in the harness
Instead of trusting a small model to route, FRIDAY adds a deterministic router in front of it: regex intent matching, fuzzy + embedding similarity, and explicit confidence bands. Known phrasings hit a real tool with certainty. The model never gets to invent a result for an action it didn't perform.
A deterministic router, cheapest layer first.
Every request walks an ordered chain and returns at the first layer that produces a confident plan. The common things you say resolve instantly with regex — no model call, no latency, no chance to hallucinate. Only the genuinely novel falls through to the on-device planner.
- Regex intent matchingFirst-match parsers turn known phrasings straight into a tool call at confidence 1.0.
- Similarity, not guessworkFuzzy (rapidfuzz) and embedding (cosine) layers catch STT slips and paraphrases.
- Confidence bandsNear-misses ask “did you mean…?” instead of silently picking the wrong tool.
- Learns your phrasingA wording you confirm a few times is promoted to deterministic dispatch.
Deterministic regex parsers — first match wins
Alias / pattern / context-term score (min 80)
rapidfuzz token-set ratio — catches STT slips & typos
Cosine similarity over capability embeddings
On-device LLM synthesises a tool plan
Conversational reply when nothing routes
A regex hit at L1 short-circuits the entire stack — the common phrasings never wake the model, and the model never gets the chance to invent a result.
The honest tradeoff
Determinism is a deal, not magic. A regex router only fires for phrasings it was taught — give FRIDAY a complex command it has never seen and the deterministic layers won't catch it; it falls back to the local planner, which is far weaker than a frontier model. This is the price of local-first: every tool we build, we also teach FRIDAY the words that drive it. The upside is reliability, privacy, and zero cloud dependence for everything it does know — and it learns new phrasings as you use it.
Privacy isn't a setting. It's the architecture.
Most assistants ship your voice and your context to someone else's servers. FRIDAY inverts that: the machine in front of you is the whole system. The network is an opt-in peripheral, not the brain.
Nothing leaves by default
STT, chat, planner, vision, and embeddings are all local GGUF / ONNX models. No account, no cloud inference, no telemetry.
Online is opt-in & consented
Web search and browser automation exist — but a consent gate asks before any capability reaches the network, and it's logged.
Your memory, on your disk
A three-tier memory (episodic, semantic, procedural) lives in local SQLite + a Chroma vector index. It's yours to inspect, export, or wipe.
Runs on modest hardware
8 GB RAM gets you going; CUDA is auto-used when present. Quantized models keep the whole stack on a laptop.
One assistant. Twenty-eight capabilities.
Voice I/O
“Hey Friday” wake word, faster-whisper STT, Piper neural TTS with barge-in. Falls back to text chat.
Natural conversation
Local chat model with session-aware turns, custom personas, and three-tier memory.
System control
Brightness, volume, screen lock/unlock, screenshots, app launch, window queries, clipboard.
Document intelligence
Index and ask questions over your PDFs, Office docs & Markdown via local RAG.
Vision (VLM)
Screenshot explainer, OCR, screen summarizer, UI-element finder, code debugger — local SmolVLM2.
Online skills (opt-in)
Browser automation, web & quick-answer search, news, world monitoring, weather.
Productivity
Reminders, calendar events, notes, tasks, goals, focus sessions, dictation.
Extensible
Add a capability plus an intent pattern; optional external MCP across 28 modules.
Privacy & safety
Ask-before-online consent, scoped security tooling (lab mode), a local audit log.
One prompt box. Three ways to drive it.
Every input surface — terminal, HUD, and the Telegram bot — understands the same prefix grammar. Talk to it, slash a capability, or drop into a real shell, all from the same line.
Runs through the deterministic router — voice or typed, same path.
Pre-routing dispatch straight to a capability — no LLM in the loop.
A PTY-backed interactive shell. venv-aware, screen-lock gated.
Pipes stdin to the running command — sudo prompts, y/n, read.
Your machine, in your pocket.
Connect a Telegram bot and FRIDAY becomes fully remote-controllable — not just chat, but the whole control surface: shell, slash commands, file Q&A, voice notes, and human-in-the-loop approvals for anything sensitive.
Bring the assistant home.
Clone it, run one setup script, and say hello. Linux and Windows, MIT licensed, no sign-up.