Files
localchat/README.md
T
bdeb1337 d1b512ff77 feat: localchat, a Go + Templ + HTMX chat app for local Gemma models
Streams replies over SSE from any OpenAI-compatible server (oMLX,
llama.cpp, Ollama, ...). Single binary that can install itself as an
OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
2026-10-05 07:24:09 +02:00

6.0 KiB
Raw Blame History

localchat

A small, private chat app for a local, open-weight AI model: Go + Templ + HTMX on top of any OpenAI-compatible model server. It defaults to Gemma 4 E2B, which runs on a normal laptop.

  • 🔒 Private: prompts and replies never leave your machine.
  • ✈️ Works offline once the model is downloaded.
  • 💸 Free to run: no API keys, no per-token billing.
  • 🔁 Model-agnostic: oMLX, llama.cpp, Ollama, LM Studio, Lemonade, vLLM… change one env var.
  • 📦 One ~9 MB binary for Linux, macOS and Windows. It can install itself as a background service.
  • 🐳 Or docker compose up for the app plus llama.cpp plus Gemma in one go.

screenshot

Quick start

Option A: containers (easiest)

docker compose up -d        # first run downloads Gemma 4 E2B (~3 GB)
open http://localhost:3000

To use a different model, set LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d.

Option B: binary + a model server you already run

make build                  # or grab a binary from `make dist`
cp .env.example .env        # point it at your model server
./bin/localchat             # → http://127.0.0.1:3000

Common LLM_BASE_URL values:

Server LLM_BASE_URL
llama.cpp llama-server http://127.0.0.1:8080/v1
oMLX (Apple Silicon) http://127.0.0.1:8000/v1
Ollama http://127.0.0.1:11434/v1
LM Studio http://127.0.0.1:1234/v1
Lemonade (AMD) http://127.0.0.1:13305/api/v1

For example, to get a model running quickly with llama.cpp:

brew install llama.cpp      # or a release from github.com/ggml-org/llama.cpp
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj

Option C: run it as a background service

The binary registers itself with the OS service manager (systemd, launchd or the Windows Service Control Manager):

./bin/localchat install --user   # per-user (launchd agent / systemd --user); drop --user for system-wide (needs sudo/admin)
./bin/localchat start --user
./bin/localchat status --user
./bin/localchat stop --user && ./bin/localchat uninstall --user

The service gets the absolute path of your .env (or --env-file), so it runs with the same settings you tested with.

Configuration

Set these as environment variables or in a .env file. Real environment variables take precedence over the file.

Variable Default
LLM_BASE_URL http://127.0.0.1:8000/v1 OpenAI-compatible API root
LLM_MODEL (first model the server lists) model id
LLM_API_KEY bearer token, if your server needs one
SYSTEM_PROMPT friendly, concise assistant give it a personality or a purpose
LOCALCHAT_TITLE localchat name in the header and tab
LOCALCHAT_ADDR 127.0.0.1:3000 listen address (0.0.0.0:3000 to share on your LAN)
MAX_HISTORY 20 past messages sent as context
LLM_TIMEOUT 5m maximum time for one reply

How it works

Browser ── HTMX + SSE extension
  │  POST /chat              → returns the user bubble + an empty reply bubble
  │  GET  /chat/stream/{id}  ← server-sent events: "token" (append) … "done" (swap in Markdown)
Go (net/http + Templ)
  │  POST /v1/chat/completions  {stream: true}
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
  1. The form posts with hx-post. The server stores the message and returns two Templ fragments: your message, and an assistant bubble with sse-connect="/chat/stream/{id}".
  2. The SSE handler claims that reply, sends the conversation to the model, and forwards each chunk as an HTML-escaped token event. HTMX appends each one (hx-swap="beforeend"), so the reply types out live.
  3. When the model finishes, a done event replaces the whole bubble with the reply rendered as Markdown (goldmark, with raw HTML stripped). That also removes the sse-connect element, which closes the stream.
  4. Browsers automatically reconnect an EventSource. A reconnect for a reply that is already claimed or finished gets the final state instead of a second generation.

There is no JavaScript framework and no build step. The only JS is htmx, its SSE extension and about 40 lines of UX glue. All of it is embedded in the binary, so the app works fully offline.

Project layout

cmd/localchat/          entry point, CLI + service install/start/stop
internal/config/        env + .env loading
internal/llm/           tiny OpenAI-compatible streaming client (stdlib only)
internal/chat/          in-memory conversations, one per browser session
internal/web/           routes, SSE streaming, embedded static assets
internal/web/views/     Templ components (*.templ → generated *_templ.go)
scripts/e2e.py          Playwright browser test against a real model
compose.yaml, Dockerfile

Development

make test     # unit + handler tests with a fake model server, -race
make run      # build + run with ./.env
make e2e      # headless browser test against the running app + real model
make dist     # cross-compile linux/darwin/windows × amd64/arm64

After editing a .templ file, run make generate (or go tool templ generate --watch). The generated *_templ.go files are committed, so go build and Docker builds don't need Templ installed.

Good to know

  • Conversations live in memory and are lost on restart. That's on purpose for privacy, and simple enough to swap for SQLite.
  • There is no authentication. It listens on localhost by default. Put a reverse proxy with auth in front before exposing it.
  • Speed in Docker depends on how many CPUs the Docker VM gets. With Docker Desktop's 1-CPU setting, Gemma E2B takes about 20 s to start replying, so give it more cores. On a Mac, running the model natively (oMLX or llama-server, which use the GPU via Metal) is far faster: about 1.5 s to the first token on an M3.
  • GPU in Docker depends on the platform. The compose file runs on CPU everywhere, which E2B handles fine. On Linux with an NVIDIA or AMD GPU, use the server-cuda or server-rocm llama.cpp image tags.