FRIDAY
Local-first · voice-native · open source

The assistant that stays on your machine.

FRIDAY listens, reasons, and acts entirely on-device. Speech-to-text, the conversational model, planning, vision, and memory all run on your hardware — it reaches the internet only when you ask. Say "Hey Friday" and it gets to work, narrating progress out loud.

install · linux
git clone https://github.com/SanthoshReddy352/Friday_Linux.git
cd Friday_Linux && ./setup.sh
python main.py
Hey Friday — set brightness to 60What's on my calendar today?Find the file design final report/deep RISC-V vs ARM in 2026!sudo systemctl restart nginxSummarize this PDFTake a screenshot and explain it
100%
on-device reasoning
28
capability modules
6-layer
routing pipeline
13
slash commands
Telegram
+ Discord remote
0
accounts · 0 telemetry
The problem with local-first agents

Cloud agents are smart because the model is smart. Take the cloud away and the harness falls apart.

The current wave of agent harnesses — the loops that read your request, pick a tool, and fill its arguments — lean entirely on a frontier model to understand intent and route correctly. Run that same harness on a model small enough to live on your laptop and it breaks: small models miss the intent, call the wrong tool, or worse —

Small models hallucinate success

Ask a 0.8B chat model to set your brightness and — with no tool actually wired — it will cheerfully reply "Brightness set to 60." Nothing happened. The model fabricated a plausible-sounding result because that's what language models do when they can't really act.

FRIDAY bridges the gap in the harness

Instead of trusting a small model to route, FRIDAY adds a deterministic router in front of it: regex intent matching, fuzzy + embedding similarity, and explicit confidence bands. Known phrasings hit a real tool with certainty. The model never gets to invent a result for an action it didn't perform.

How FRIDAY routes

A deterministic router, cheapest layer first.

Every request walks an ordered chain and returns at the first layer that produces a confident plan. The common things you say resolve instantly with regex — no model call, no latency, no chance to hallucinate. Only the genuinely novel falls through to the on-device planner.

  • Regex intent matching
    First-match parsers turn known phrasings straight into a tool call at confidence 1.0.
  • Similarity, not guesswork
    Fuzzy (rapidfuzz) and embedding (cosine) layers catch STT slips and paraphrases.
  • Confidence bands
    Near-misses ask “did you mean…?” instead of silently picking the wrong tool.
  • Learns your phrasing
    A wording you confirm a few times is promoted to deterministic dispatch.
Read how routing works
voice / text in
cheapest layer that resolves wins →
L1
Intent Recognizerconfidence 1.0

Deterministic regex parsers — first match wins

L2
Route Scorerscore ≥ 80

Alias / pattern / context-term score (min 80)

L2b
Lexical Routerfuzzy ≥ 88

rapidfuzz token-set ratio — catches STT slips & typos

L3
Embedding Routercosine band + confirm

Cosine similarity over capability embeddings

L4
Local Planner (4B)generative

On-device LLM synthesises a tool plan

L5
Chat Fallbackchat

Conversational reply when nothing routes

A regex hit at L1 short-circuits the entire stack — the common phrasings never wake the model, and the model never gets the chance to invent a result.

The honest tradeoff

Determinism is a deal, not magic. A regex router only fires for phrasings it was taught — give FRIDAY a complex command it has never seen and the deterministic layers won't catch it; it falls back to the local planner, which is far weaker than a frontier model. This is the price of local-first: every tool we build, we also teach FRIDAY the words that drive it. The upside is reliability, privacy, and zero cloud dependence for everything it does know — and it learns new phrasings as you use it.

Local-first, in full

Privacy isn't a setting. It's the architecture.

Most assistants ship your voice and your context to someone else's servers. FRIDAY inverts that: the machine in front of you is the whole system. The network is an opt-in peripheral, not the brain.

Nothing leaves by default

STT, chat, planner, vision, and embeddings are all local GGUF / ONNX models. No account, no cloud inference, no telemetry.

Online is opt-in & consented

Web search and browser automation exist — but a consent gate asks before any capability reaches the network, and it's logged.

Your memory, on your disk

A three-tier memory (episodic, semantic, procedural) lives in local SQLite + a Chroma vector index. It's yours to inspect, export, or wipe.

Runs on modest hardware

8 GB RAM gets you going; CUDA is auto-used when present. Quantized models keep the whole stack on a laptop.

What it can do

One assistant. Twenty-eight capabilities.

Full capability reference
voice

Voice I/O

“Hey Friday” wake word, faster-whisper STT, Piper neural TTS with barge-in. Falls back to text chat.

chat

Natural conversation

Local chat model with session-aware turns, custom personas, and three-tier memory.

system

System control

Brightness, volume, screen lock/unlock, screenshots, app launch, window queries, clipboard.

docs

Document intelligence

Index and ask questions over your PDFs, Office docs & Markdown via local RAG.

vision

Vision (VLM)

Screenshot explainer, OCR, screen summarizer, UI-element finder, code debugger — local SmolVLM2.

web

Online skills (opt-in)

Browser automation, web & quick-answer search, news, world monitoring, weather.

tasks

Productivity

Reminders, calendar events, notes, tasks, goals, focus sessions, dictation.

ext

Extensible

Add a capability plus an intent pattern; optional external MCP across 28 modules.

safe

Privacy & safety

Ask-before-online consent, scoped security tooling (lab mode), a local audit log.

Command it from anywhere

One prompt box. Three ways to drive it.

Every input surface — terminal, HUD, and the Telegram bot — understands the same prefix grammar. Talk to it, slash a capability, or drop into a real shell, all from the same line.

·Plain text
set brightness to 60

Runs through the deterministic router — voice or typed, same path.

/Slash command
/deep RISC-V vs ARM

Pre-routing dispatch straight to a capability — no LLM in the loop.

!Shell command
!sudo systemctl restart nginx

A PTY-backed interactive shell. venv-aware, screen-lock gated.

>Shell follow-up
> your-password

Pipes stdin to the running command — sudo prompts, y/n, read.

Telegram & Discord

Your machine, in your pocket.

Connect a Telegram bot and FRIDAY becomes fully remote-controllable — not just chat, but the whole control surface: shell, slash commands, file Q&A, voice notes, and human-in-the-loop approvals for anything sensitive.

Drive the full assistant
Message the bot and the entire turn pipeline runs on your machine, replying in the chat.
Shell from your phone
! commands and > follow-ups work remotely — restart a service, run a script, answer a sudo prompt.
13 slash commands
/new /web /quick /fast /deep /research /fetch /crawl /screenshot /voice /lock /unlock /help — with Telegram autocomplete.
Upload a document
Drop a PDF / DOCX / XLSX / MD into the chat; it loads into RAG and you ask questions about it.
Send a voice note
Transcribed by the same local Whisper STT, then processed as a normal turn.
Approval gates
Security workflows send a yes/no prompt and block until you approve — refuse-by-default if you're offline.
Proactive push
Reminders, goal check-ins, and triggers reach you on Telegram or Discord when you're away.
Live, in-place replies
A typing indicator and a 💭 bubble that morphs into the answer; long replies auto-chunked.
Set up remote control
F
FRIDAY
online · on your machine
!sudo systemctl restart nginx
21:04
[sudo] password for you:
Awaiting password — reply with > ••••
21:04
> ••••••••
21:04
✓ nginx restarted · [exit 0]
21:04
slash command
/deep RISC-V vs ARM in 2026
21:06
Researching across sources… briefing saved to disk.
21:06
voice note + approval
0:04
21:09
Run a port scan on the lab host?
approvedeny
21:09
Message FRIDAY…

Bring the assistant home.

Clone it, run one setup script, and say hello. Linux and Windows, MIT licensed, no sign-up.