Local voice assistant with Ollama
Microphone → speech recognition → local LLM → voice. A whole assistant on your own machine, nothing sent to the cloud.
In this post
A voice assistant is four blocks: listen, understand, think, speak. Each of them can run locally today — on a laptop, with no API key and no recordings leaving your machine. This post wires those blocks into a working Python prototype. Our Jarvis grew from the same trunk.
The blocks
| Stage | Tool | Why |
|---|---|---|
| Listen | sounddevice + a simple VAD | no native deps besides PortAudio |
| Understand (STT) | faster-whisper | Whisper on CPU/GPU; the small model is enough for conversation |
| Think (LLM) | Ollama | one ollama run, a local API on :11434 |
| Speak (TTS) | piper | fast, light voices in many languages |
1. Listen: when has the user stopped talking?
The simplest silence detector: record 30 ms blocks, compute energy (RMS), and call the utterance finished after ~800 ms below a threshold.
import numpy as np, sounddevice as sd
SR = 16000
def record_utterance(threshold=0.01, silence_ms=800):
chunks, silent = [], 0
with sd.InputStream(samplerate=SR, channels=1, dtype="float32") as stream:
while True:
block, _ = stream.read(int(SR * 0.03))
chunks.append(block)
rms = float(np.sqrt(np.mean(block ** 2)))
silent = silent + 30 if rms < threshold else 0
if chunks and silent >= silence_ms and len(chunks) > 20:
break
return np.concatenate(chunks).flatten()
The same RMS measurement later drives a character’s “mouth” — that is how lipsync works in Simlokatorzy.
2. Understand: faster-whisper
from faster_whisper import WhisperModel
stt = WhisperModel("small", device="cpu", compute_type="int8")
def transcribe(audio):
segments, _ = stt.transcribe(audio, language="en", vad_filter=True)
return " ".join(s.text for s in segments).strip()
The small model in int8 transcribes a sentence in about a second on an ordinary laptop. For English you can drop to base.
3. Think: Ollama
Run ollama pull llama3.1 once (or any model that fits your memory). Ollama exposes a local HTTP API; requests is all you need:
import requests
SYSTEM = "You are a concise assistant. Answer in at most two sentences."
history = [{"role": "system", "content": SYSTEM}]
def think(user_text):
history.append({"role": "user", "content": user_text})
r = requests.post("http://localhost:11434/api/chat", json={
"model": "llama3.1", "messages": history, "stream": False,
})
answer = r.json()["message"]["content"]
history.append({"role": "assistant", "content": answer})
return answer
The history list is the whole “memory” of the conversation. For longer sessions trim it to the last N messages or summarise — the model has a limited context window.
4. Speak: piper
Piper generates speech from an .onnx voice file, offline:
import subprocess
def speak(text, voice="en_US-lessac-medium.onnx"):
subprocess.run(
["piper", "--model", voice, "--output-raw"],
input=text.encode(), stdout=subprocess.PIPE, check=True,
)
The raw PCM from --output-raw can be played straight through sounddevice without writing a file — lower latency.
The loop
while True:
audio = record_utterance()
text = transcribe(audio)
if not text:
continue
print("You:", text)
reply = think(text)
print("Assistant:", reply)
speak(reply)
The full cycle (end of speech to start of the reply) on a laptop without a GPU is 3–5 seconds. The LLM is the biggest cost; a smaller model, or streaming Ollama’s answer ("stream": True and feeding sentences to TTS as they arrive) shortens it noticeably.
Where to go next
- Wake word: only start
record_utterance()after “hey, Jarvis” (e.g.openwakeword). - Tools: instead of a plain answer, ask the model for JSON
{"action": "...", "args": {...}}and execute it — the first step from chat to agent. - Barge-in: keep listening while speaking and cut the TTS when the user interrupts.
Everything above runs offline. Your recordings never leave the computer — a rare luxury these days.