Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
6.0 KiB
localchat
A small, private chat app for a local, open-weight AI model: Go + Templ + HTMX on top of any OpenAI-compatible model server. It defaults to Gemma 4 E2B, which runs on a normal laptop.
- 🔒 Private: prompts and replies never leave your machine.
- ✈️ Works offline once the model is downloaded.
- 💸 Free to run: no API keys, no per-token billing.
- 🔁 Model-agnostic: oMLX, llama.cpp, Ollama, LM Studio, Lemonade, vLLM… change one env var.
- 📦 One ~9 MB binary for Linux, macOS and Windows. It can install itself as a background service.
- 🐳 Or
docker compose upfor the app plus llama.cpp plus Gemma in one go.
Quick start
Option A: containers (easiest)
docker compose up -d # first run downloads Gemma 4 E2B (~3 GB)
open http://localhost:3000
To use a different model, set LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d.
Option B: binary + a model server you already run
make build # or grab a binary from `make dist`
cp .env.example .env # point it at your model server
./bin/localchat # → http://127.0.0.1:3000
Common LLM_BASE_URL values:
| Server | LLM_BASE_URL |
|---|---|
llama.cpp llama-server |
http://127.0.0.1:8080/v1 |
| oMLX (Apple Silicon) | http://127.0.0.1:8000/v1 |
| Ollama | http://127.0.0.1:11434/v1 |
| LM Studio | http://127.0.0.1:1234/v1 |
| Lemonade (AMD) | http://127.0.0.1:13305/api/v1 |
For example, to get a model running quickly with llama.cpp:
brew install llama.cpp # or a release from github.com/ggml-org/llama.cpp
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj
Option C: run it as a background service
The binary registers itself with the OS service manager (systemd, launchd or the Windows Service Control Manager):
./bin/localchat install --user # per-user (launchd agent / systemd --user); drop --user for system-wide (needs sudo/admin)
./bin/localchat start --user
./bin/localchat status --user
./bin/localchat stop --user && ./bin/localchat uninstall --user
The service gets the absolute path of your .env (or --env-file), so it runs with the same settings you tested with.
Configuration
Set these as environment variables or in a .env file. Real environment variables take precedence over the file.
| Variable | Default | |
|---|---|---|
LLM_BASE_URL |
http://127.0.0.1:8000/v1 |
OpenAI-compatible API root |
LLM_MODEL |
(first model the server lists) | model id |
LLM_API_KEY |
bearer token, if your server needs one | |
SYSTEM_PROMPT |
friendly, concise assistant | give it a personality or a purpose |
LOCALCHAT_TITLE |
localchat |
name in the header and tab |
LOCALCHAT_ADDR |
127.0.0.1:3000 |
listen address (0.0.0.0:3000 to share on your LAN) |
MAX_HISTORY |
20 |
past messages sent as context |
LLM_TIMEOUT |
5m |
maximum time for one reply |
How it works
Browser ── HTMX + SSE extension
│ POST /chat → returns the user bubble + an empty reply bubble
│ GET /chat/stream/{id} ← server-sent events: "token" (append) … "done" (swap in Markdown)
Go (net/http + Templ)
│ POST /v1/chat/completions {stream: true}
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
- The form posts with
hx-post. The server stores the message and returns two Templ fragments: your message, and an assistant bubble withsse-connect="/chat/stream/{id}". - The SSE handler claims that reply, sends the conversation to the model, and forwards each chunk as an HTML-escaped
tokenevent. HTMX appends each one (hx-swap="beforeend"), so the reply types out live. - When the model finishes, a
doneevent replaces the whole bubble with the reply rendered as Markdown (goldmark, with raw HTML stripped). That also removes thesse-connectelement, which closes the stream. - Browsers automatically reconnect an EventSource. A reconnect for a reply that is already claimed or finished gets the final state instead of a second generation.
There is no JavaScript framework and no build step. The only JS is htmx, its SSE extension and about 40 lines of UX glue. All of it is embedded in the binary, so the app works fully offline.
Project layout
cmd/localchat/ entry point, CLI + service install/start/stop
internal/config/ env + .env loading
internal/llm/ tiny OpenAI-compatible streaming client (stdlib only)
internal/chat/ in-memory conversations, one per browser session
internal/web/ routes, SSE streaming, embedded static assets
internal/web/views/ Templ components (*.templ → generated *_templ.go)
scripts/e2e.py Playwright browser test against a real model
compose.yaml, Dockerfile
Development
make test # unit + handler tests with a fake model server, -race
make run # build + run with ./.env
make e2e # headless browser test against the running app + real model
make dist # cross-compile linux/darwin/windows × amd64/arm64
After editing a .templ file, run make generate (or go tool templ generate --watch). The generated *_templ.go files are committed, so go build and Docker builds don't need Templ installed.
Good to know
- Conversations live in memory and are lost on restart. That's on purpose for privacy, and simple enough to swap for SQLite.
- There is no authentication. It listens on localhost by default. Put a reverse proxy with auth in front before exposing it.
- Speed in Docker depends on how many CPUs the Docker VM gets. With Docker Desktop's 1-CPU setting, Gemma E2B takes about 20 s to start replying, so give it more cores. On a Mac, running the model natively (oMLX or
llama-server, which use the GPU via Metal) is far faster: about 1.5 s to the first token on an M3. - GPU in Docker depends on the platform. The compose file runs on CPU everywhere, which E2B handles fine. On Linux with an NVIDIA or AMD GPU, use the
server-cudaorserver-rocmllama.cpp image tags.
