Files
bdeb1337 bb71b49fde feat: localchat, a Go + Templ + HTMX chat app for local Gemma models
Streams replies over SSE from any OpenAI-compatible server (oMLX,
llama.cpp, Ollama, ...). Single binary that can install itself as an
OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
2026-10-05 07:25:42 +02:00

6.2 KiB
Raw Permalink Blame History

localchat

localchat gives you a browser chat UI for an open-weight AI model running on your own machine. It's a Go + Templ + HTMX app that talks to any OpenAI-compatible model server, and it defaults to Gemma 4 E2B, which runs on a normal laptop.

  • 🔒 Your prompts and replies stay on your machine, and the app works offline once you have the model.
  • 💸 You need no API key and pay nothing per token.
  • 🔁 You can point it at oMLX, llama.cpp, Ollama, LM Studio, Lemonade or vLLM by changing one environment variable.
  • 📦 You get a ~9 MB binary for Linux, macOS or Windows that can install itself as a background service. docker compose up starts the app, llama.cpp and Gemma together.

screenshot

Quick start

Option A: containers

docker compose up -d        # first run downloads Gemma 4 E2B (~3 GB)
open http://localhost:3000

To use a different model, set LLM_HF_REPO:

LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d

Option B: binary + your own model server

make build                  # or grab a binary from `make dist`
cp .env.example .env        # point it at your model server
./bin/localchat             # → http://127.0.0.1:3000

Set LLM_BASE_URL to match your server:

Server LLM_BASE_URL
llama.cpp llama-server http://127.0.0.1:8080/v1
oMLX (Apple Silicon) http://127.0.0.1:8000/v1
Ollama http://127.0.0.1:11434/v1
LM Studio http://127.0.0.1:1234/v1
Lemonade (AMD) http://127.0.0.1:13305/api/v1

If you don't run a model server yet, llama.cpp gets you one in two commands:

brew install llama.cpp      # or a release from github.com/ggml-org/llama.cpp
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj

Option C: background service

You can register localchat with your OS service manager: systemd on Linux, launchd on macOS, or the Service Control Manager on Windows.

./bin/localchat install --user   # per-user service; drop --user for a system-wide one (needs sudo/admin)
./bin/localchat start --user
./bin/localchat status --user
./bin/localchat stop --user && ./bin/localchat uninstall --user

install writes the absolute path of your .env (or --env-file) into the service definition, so the service runs with the settings you tested.

Configuration

Set these as environment variables or in a .env file. If you set both, the environment variable wins.

Variable Default
LLM_BASE_URL http://127.0.0.1:8000/v1 OpenAI-compatible API root
LLM_MODEL (first model the server lists) model id
LLM_API_KEY bearer token, if your server needs one
SYSTEM_PROMPT friendly, concise assistant give it a personality or a purpose
LOCALCHAT_TITLE localchat name in the header and tab
LOCALCHAT_ADDR 127.0.0.1:3000 listen address (0.0.0.0:3000 to share on your LAN)
MAX_HISTORY 20 past messages sent as context
LLM_TIMEOUT 5m time limit for one reply

Architecture

Browser ── HTMX + SSE extension
  │  POST /chat              → returns the user bubble + an empty reply bubble
  │  GET  /chat/stream/{id}  ← server-sent events: "token" (append) … "done" (swap in Markdown)
Go (net/http + Templ)
  │  POST /v1/chat/completions  {stream: true}
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
  1. Your browser posts the form with hx-post. The server stores the message and returns two Templ fragments: your message and an assistant bubble with sse-connect="/chat/stream/{id}".
  2. The SSE handler claims the reply, sends the conversation to the model and forwards each chunk as an HTML-escaped token event. HTMX appends the chunks with hx-swap="beforeend", so you see the reply appear as the model writes it.
  3. Once the model finishes, the handler sends a done event. HTMX swaps the bubble for the reply rendered as Markdown (goldmark, raw HTML stripped), and removing the sse-connect element closes the stream.
  4. Browsers reconnect a dropped EventSource on their own. If a browser reconnects for a reply that another request already claimed or finished, the handler returns the final state and leaves the model alone.

You won't find a JavaScript framework or a frontend build step. The JavaScript comes down to htmx, its SSE extension and about 40 lines of UX code, and localchat embeds all of it in the binary so the app runs offline.

Project layout

cmd/localchat/          entry point, CLI + service install/start/stop
internal/config/        env + .env loading
internal/llm/           OpenAI-compatible streaming client (standard library)
internal/chat/          in-memory conversations, one per browser session
internal/web/           routes, SSE streaming, embedded static assets
internal/web/views/     Templ components (*.templ → generated *_templ.go)
scripts/e2e.py          Playwright browser test against a real model
compose.yaml, Dockerfile

Development

make test     # unit + handler tests with a fake model server, -race
make run      # build + run with ./.env
make e2e      # headless browser test against the running app + real model
make dist     # cross-compile linux/darwin/windows × amd64/arm64

After you edit a .templ file, run make generate (or go tool templ generate --watch). The repo includes the generated *_templ.go files, so you can run go build or a Docker build without Templ installed.

Good to know

  • localchat keeps conversations in memory, so a restart clears them. That keeps chats private; add SQLite if you want history to survive a restart.
  • localchat has no authentication and listens on localhost by default. Put a reverse proxy with auth in front before you expose it.
  • In Docker, speed depends on how many CPUs you give the Docker VM. With Docker Desktop set to 1 CPU, Gemma E2B needs about 20 s before it starts replying. On a Mac, a native model server (oMLX or llama-server, both using the GPU through Metal) sends the first token in about a second on an M3.
  • The compose file runs the model on the CPU, which works on any platform and suits E2B. On Linux with an NVIDIA or AMD GPU, switch to the server-cuda or server-rocm llama.cpp image.