feat: localchat, a Go + Templ + HTMX chat app for local Gemma models

Streams replies over SSE from any OpenAI-compatible server (oMLX,
llama.cpp, Ollama, ...). Single binary that can install itself as an
OS service; docker compose bundles llama.cpp + Gemma 4 E2B.
This commit is contained in:
bdeb1337 committed 2026-10-05 07:24:09 +02:00
commit d1b512ff77
28 files changed
+2457

No files matched your search

+125
View File
@@ -0,0 +1,125 @@
# localchat
A small, private chat app for a **local, open-weight AI model**: **Go + Templ + HTMX** on top of any OpenAI-compatible model server. It defaults to **Gemma 4 E2B**, which runs on a normal laptop.
- 🔒 **Private:** prompts and replies never leave your machine.
- ✈️ **Works offline** once the model is downloaded.
- 💸 **Free to run:** no API keys, no per-token billing.
- 🔁 **Model-agnostic:** oMLX, llama.cpp, Ollama, LM Studio, Lemonade, vLLM… change one env var.
- 📦 **One ~9 MB binary** for Linux, macOS and Windows. It can install itself as a background service.
- 🐳 **Or `docker compose up`** for the app plus llama.cpp plus Gemma in one go.
![screenshot](docs/screenshot.png)
## Quick start
### Option A: containers (easiest)
```sh
docker compose up -d # first run downloads Gemma 4 E2B (~3 GB)
open http://localhost:3000
```
To use a different model, set `LLM_HF_REPO=unsloth/gemma-4-E4B-it-GGUF:Q4_K_M docker compose up -d`.
### Option B: binary + a model server you already run
```sh
make build # or grab a binary from `make dist`
cp .env.example .env # point it at your model server
./bin/localchat # → http://127.0.0.1:3000
```
Common `LLM_BASE_URL` values:
| Server | `LLM_BASE_URL` |
|---|---|
| llama.cpp `llama-server` | `http://127.0.0.1:8080/v1` |
| oMLX (Apple Silicon) | `http://127.0.0.1:8000/v1` |
| Ollama | `http://127.0.0.1:11434/v1` |
| LM Studio | `http://127.0.0.1:1234/v1` |
| Lemonade (AMD) | `http://127.0.0.1:13305/api/v1` |
For example, to get a model running quickly with llama.cpp:
```sh
brew install llama.cpp # or a release from github.com/ggml-org/llama.cpp
llama-server -hf unsloth/gemma-4-E2B-it-GGUF:Q4_K_M --no-mmproj
```
### Option C: run it as a background service
The binary registers itself with the OS service manager (systemd, launchd or the Windows Service Control Manager):
```sh
./bin/localchat install --user # per-user (launchd agent / systemd --user); drop --user for system-wide (needs sudo/admin)
./bin/localchat start --user
./bin/localchat status --user
./bin/localchat stop --user && ./bin/localchat uninstall --user
```
The service gets the absolute path of your `.env` (or `--env-file`), so it runs with the same settings you tested with.
## Configuration
Set these as environment variables or in a `.env` file. Real environment variables take precedence over the file.
| Variable | Default | |
|---|---|---|
| `LLM_BASE_URL` | `http://127.0.0.1:8000/v1` | OpenAI-compatible API root |
| `LLM_MODEL` | *(first model the server lists)* | model id |
| `LLM_API_KEY` | | bearer token, if your server needs one |
| `SYSTEM_PROMPT` | friendly, concise assistant | give it a personality or a purpose |
| `LOCALCHAT_TITLE` | `localchat` | name in the header and tab |
| `LOCALCHAT_ADDR` | `127.0.0.1:3000` | listen address (`0.0.0.0:3000` to share on your LAN) |
| `MAX_HISTORY` | `20` | past messages sent as context |
| `LLM_TIMEOUT` | `5m` | maximum time for one reply |
## How it works
```
Browser ── HTMX + SSE extension
│ POST /chat → returns the user bubble + an empty reply bubble
│ GET /chat/stream/{id} ← server-sent events: "token" (append) … "done" (swap in Markdown)
Go (net/http + Templ)
│ POST /v1/chat/completions {stream: true}
Model server (llama.cpp / oMLX / Ollama / …) ── Gemma 4 E2B
```
1. The form posts with `hx-post`. The server stores the message and returns two Templ fragments: your message, and an assistant bubble with `sse-connect="/chat/stream/{id}"`.
2. The SSE handler claims that reply, sends the conversation to the model, and forwards each chunk as an HTML-escaped `token` event. HTMX appends each one (`hx-swap="beforeend"`), so the reply types out live.
3. When the model finishes, a `done` event replaces the whole bubble with the reply rendered as Markdown (goldmark, with raw HTML stripped). That also removes the `sse-connect` element, which closes the stream.
4. Browsers automatically reconnect an EventSource. A reconnect for a reply that is already claimed or finished gets the final state instead of a second generation.
There is no JavaScript framework and no build step. The only JS is htmx, its SSE extension and about 40 lines of UX glue. All of it is embedded in the binary, so the app works fully offline.
## Project layout
```
cmd/localchat/ entry point, CLI + service install/start/stop
internal/config/ env + .env loading
internal/llm/ tiny OpenAI-compatible streaming client (stdlib only)
internal/chat/ in-memory conversations, one per browser session
internal/web/ routes, SSE streaming, embedded static assets
internal/web/views/ Templ components (*.templ → generated *_templ.go)
scripts/e2e.py Playwright browser test against a real model
compose.yaml, Dockerfile
```
## Development
```sh
make test # unit + handler tests with a fake model server, -race
make run # build + run with ./.env
make e2e # headless browser test against the running app + real model
make dist # cross-compile linux/darwin/windows × amd64/arm64
```
After editing a `.templ` file, run `make generate` (or `go tool templ generate --watch`). The generated `*_templ.go` files are committed, so `go build` and Docker builds don't need Templ installed.
## Good to know
- Conversations live in memory and are lost on restart. That's on purpose for privacy, and simple enough to swap for SQLite.
- There is no authentication. It listens on localhost by default. Put a reverse proxy with auth in front before exposing it.
- Speed in Docker depends on how many CPUs the Docker VM gets. With Docker Desktop's 1-CPU setting, Gemma E2B takes about 20 s to start replying, so give it more cores. On a Mac, running the model natively (oMLX or `llama-server`, which use the GPU via Metal) is far faster: about 1.5 s to the first token on an M3.
- GPU in Docker depends on the platform. The compose file runs on CPU everywhere, which E2B handles fine. On Linux with an NVIDIA or AMD GPU, use the `server-cuda` or `server-rocm` llama.cpp image tags.