← Blog

Local voice assistant with Ollama

Microphone → speech recognition → local LLM → voice. A whole assistant on your own machine, nothing sent to the cloud.

Kacper Walczak · 6 March 2026 · 3 min read

In this post
  1. The blocks
  2. 1. Listen: when has the user stopped talking?
  3. 2. Understand: faster-whisper
  4. 3. Think: Ollama
  5. 4. Speak: piper
  6. The loop
  7. Where to go next

A voice assistant is four blocks: listen, understand, think, speak. Each of them can run locally today — on a laptop, with no API key and no recordings leaving your machine. This post wires those blocks into a working Python prototype. Our Jarvis grew from the same trunk.

The blocks

StageToolWhy
Listensounddevice + a simple VADno native deps besides PortAudio
Understand (STT)faster-whisperWhisper on CPU/GPU; the small model is enough for conversation
Think (LLM)Ollamaone ollama run, a local API on :11434
Speak (TTS)piperfast, light voices in many languages

1. Listen: when has the user stopped talking?

The simplest silence detector: record 30 ms blocks, compute energy (RMS), and call the utterance finished after ~800 ms below a threshold.

import numpy as np, sounddevice as sd

SR = 16000
def record_utterance(threshold=0.01, silence_ms=800):
    chunks, silent = [], 0
    with sd.InputStream(samplerate=SR, channels=1, dtype="float32") as stream:
        while True:
            block, _ = stream.read(int(SR * 0.03))
            chunks.append(block)
            rms = float(np.sqrt(np.mean(block ** 2)))
            silent = silent + 30 if rms < threshold else 0
            if chunks and silent >= silence_ms and len(chunks) > 20:
                break
    return np.concatenate(chunks).flatten()

The same RMS measurement later drives a character’s “mouth” — that is how lipsync works in Simlokatorzy.

2. Understand: faster-whisper

from faster_whisper import WhisperModel
stt = WhisperModel("small", device="cpu", compute_type="int8")

def transcribe(audio):
    segments, _ = stt.transcribe(audio, language="en", vad_filter=True)
    return " ".join(s.text for s in segments).strip()

The small model in int8 transcribes a sentence in about a second on an ordinary laptop. For English you can drop to base.

3. Think: Ollama

Run ollama pull llama3.1 once (or any model that fits your memory). Ollama exposes a local HTTP API; requests is all you need:

import requests

SYSTEM = "You are a concise assistant. Answer in at most two sentences."
history = [{"role": "system", "content": SYSTEM}]

def think(user_text):
    history.append({"role": "user", "content": user_text})
    r = requests.post("http://localhost:11434/api/chat", json={
        "model": "llama3.1", "messages": history, "stream": False,
    })
    answer = r.json()["message"]["content"]
    history.append({"role": "assistant", "content": answer})
    return answer

The history list is the whole “memory” of the conversation. For longer sessions trim it to the last N messages or summarise — the model has a limited context window.

4. Speak: piper

Piper generates speech from an .onnx voice file, offline:

import subprocess

def speak(text, voice="en_US-lessac-medium.onnx"):
    subprocess.run(
        ["piper", "--model", voice, "--output-raw"],
        input=text.encode(), stdout=subprocess.PIPE, check=True,
    )

The raw PCM from --output-raw can be played straight through sounddevice without writing a file — lower latency.

The loop

while True:
    audio = record_utterance()
    text = transcribe(audio)
    if not text:
        continue
    print("You:", text)
    reply = think(text)
    print("Assistant:", reply)
    speak(reply)

The full cycle (end of speech to start of the reply) on a laptop without a GPU is 3–5 seconds. The LLM is the biggest cost; a smaller model, or streaming Ollama’s answer ("stream": True and feeding sentences to TTS as they arrive) shortens it noticeably.

Where to go next

  • Wake word: only start record_utterance() after “hey, Jarvis” (e.g. openwakeword).
  • Tools: instead of a plain answer, ask the model for JSON {"action": "...", "args": {...}} and execute it — the first step from chat to agent.
  • Barge-in: keep listening while speaking and cut the TTS when the user interrupts.

Everything above runs offline. Your recordings never leave the computer — a rare luxury these days.