Streams replies over SSE from any OpenAI-compatible server (oMLX, llama.cpp, Ollama, ...). Single binary that can install itself as an OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
6.2 KiB
localchat
localchat gives you a browser chat UI for an open-weight AI model running on your own machine. It's a Go + Templ + HTMX app that talks to any OpenAI-compatible model server, and it defaults to Gemma 4 E2B, which runs on a normal laptop.
- 🔒 Your prompts and replies stay on your machine, and the app works offline once you have the model.
- 💸 You need no API key and pay nothing per token.
- 🔁 You can point it at oMLX, llama.cpp, Ollama, LM Studio, Lemonade or vLLM by changing one environment variable.
- 📦 You get a ~9 MB binary for Linux, macOS or Windows that can install itself as a background service.
docker compose upstarts the app, llama.cpp and Gemma together.
Quick start
Option A: containers
docker compose up -d # first run downloads Gemma 4 E2B (~3 GB)
open http://localhost:3000
To use a different model, set LLM_HF_REPO:
LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d
Option B: binary + your own model server
make build # or grab a binary from `make dist`
cp .env.example .env # point it at your model server
./bin/localchat # → http://127.0.0.1:3000
Set LLM_BASE_URL to match your server:
| Server | LLM_BASE_URL |
|---|---|
llama.cpp llama-server |
http://127.0.0.1:8080/v1 |
| oMLX (Apple Silicon) | http://127.0.0.1:8000/v1 |
| Ollama | http://127.0.0.1:11434/v1 |
| LM Studio | http://127.0.0.1:1234/v1 |
| Lemonade (AMD) | http://127.0.0.1:13305/api/v1 |
If you don't run a model server yet, llama.cpp gets you one in two commands:
brew install llama.cpp # or a release from github.com/ggml-org/llama.cpp
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj
Option C: background service
You can register localchat with your OS service manager: systemd on Linux, launchd on macOS, or the Service Control Manager on Windows.
./bin/localchat install --user # per-user service; drop --user for a system-wide one (needs sudo/admin)
./bin/localchat start --user
./bin/localchat status --user
./bin/localchat stop --user && ./bin/localchat uninstall --user
install writes the absolute path of your .env (or --env-file) into the service definition, so the service runs with the settings you tested.
Configuration
Set these as environment variables or in a .env file. If you set both, the environment variable wins.
| Variable | Default | |
|---|---|---|
LLM_BASE_URL |
http://127.0.0.1:8000/v1 |
OpenAI-compatible API root |
LLM_MODEL |
(first model the server lists) | model id |
LLM_API_KEY |
bearer token, if your server needs one | |
SYSTEM_PROMPT |
friendly, concise assistant | give it a personality or a purpose |
LOCALCHAT_TITLE |
localchat |
name in the header and tab |
LOCALCHAT_ADDR |
127.0.0.1:3000 |
listen address (0.0.0.0:3000 to share on your LAN) |
MAX_HISTORY |
20 |
past messages sent as context |
LLM_TIMEOUT |
5m |
time limit for one reply |
Architecture
Browser ── HTMX + SSE extension
│ POST /chat → returns the user bubble + an empty reply bubble
│ GET /chat/stream/{id} ← server-sent events: "token" (append) … "done" (swap in Markdown)
Go (net/http + Templ)
│ POST /v1/chat/completions {stream: true}
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
- Your browser posts the form with
hx-post. The server stores the message and returns two Templ fragments: your message and an assistant bubble withsse-connect="/chat/stream/{id}". - The SSE handler claims the reply, sends the conversation to the model and forwards each chunk as an HTML-escaped
tokenevent. HTMX appends the chunks withhx-swap="beforeend", so you see the reply appear as the model writes it. - Once the model finishes, the handler sends a
doneevent. HTMX swaps the bubble for the reply rendered as Markdown (goldmark, raw HTML stripped), and removing thesse-connectelement closes the stream. - Browsers reconnect a dropped EventSource on their own. If a browser reconnects for a reply that another request already claimed or finished, the handler returns the final state and leaves the model alone.
You won't find a JavaScript framework or a frontend build step. The JavaScript comes down to htmx, its SSE extension and about 40 lines of UX code, and localchat embeds all of it in the binary so the app runs offline.
Project layout
cmd/localchat/ entry point, CLI + service install/start/stop
internal/config/ env + .env loading
internal/llm/ OpenAI-compatible streaming client (standard library)
internal/chat/ in-memory conversations, one per browser session
internal/web/ routes, SSE streaming, embedded static assets
internal/web/views/ Templ components (*.templ → generated *_templ.go)
scripts/e2e.py Playwright browser test against a real model
compose.yaml, Dockerfile
Development
make test # unit + handler tests with a fake model server, -race
make run # build + run with ./.env
make e2e # headless browser test against the running app + real model
make dist # cross-compile linux/darwin/windows × amd64/arm64
After you edit a .templ file, run make generate (or go tool templ generate --watch). The repo includes the generated *_templ.go files, so you can run go build or a Docker build without Templ installed.
Good to know
- localchat keeps conversations in memory, so a restart clears them. That keeps chats private; add SQLite if you want history to survive a restart.
- localchat has no authentication and listens on localhost by default. Put a reverse proxy with auth in front before you expose it.
- In Docker, speed depends on how many CPUs you give the Docker VM. With Docker Desktop set to 1 CPU, Gemma E2B needs about 20 s before it starts replying. On a Mac, a native model server (oMLX or
llama-server, both using the GPU through Metal) sends the first token in about a second on an M3. - The compose file runs the model on the CPU, which works on any platform and suits E2B. On Linux with an NVIDIA or AMD GPU, switch to the
server-cudaorserver-rocmllama.cpp image.
